Security News

Cybersecurity news aggregator

MEDIUM Attacks Dark Reading

Hacker Turns AI Jailbreaks Into Offensive Attack Platform

  • What: Hacker creates offensive attack platform from AI jailbreaks
  • Impact: Russian-speaking actor 'Trim' uses compromised models for cybercrime
Read Full Article →

Informa TechTarget | SearchSecurity Cybersecurity Dive InformationWeek Channel Dive Explore our brands Dark Reading Resource Library Black Hat News Omdia Cybersecurity Advertise NEWSLETTER SIGN-UP Cybersecurity Topics World The Edge DR Technology Events Resources CYBER RISK THREAT INTELLIGENCE VULNERABILITIES & THREATS REMOTE WORKFORCE NEWS Hacker Turns AI Jailbreaks Into Offensive Attack Platform A Russian-speaking actor, "Trim," dismantled publicly available frontier models and integrated them with offensive security tools. Elizabeth Montalbano,Contributing Writer July 21, 2026 5 Min Read SOURCE: TECHA TUNGATEJA VIA ALAMY STOCK PHOTO An enterprising Russian-speaking hacker spent part of the year jailbreaking publicly available, frontier large language models (LLMs) to create a for-fee offensive cybercriminal tool, researchers have found. The cybercriminal known as "Trim" started his operation by publishing jailbreaking techniques on an underground forum in late March. He eventually turned them "into a fully productized, commercially marketed AI-powered penetration-testing platform," researchers from Cato Networks' Cato CTRL revealed in a report published today. LOADING... The operation demonstrates once again how artificial intelligence (AI) models themselves have become part of the attack surface for cybercriminals, according to the report. Moreover, other bad actors are beginning to follow this blueprint for weaponizing AI models, which increases the security risk. "Trim didn't need a vulnerability to exploit," according to the report. "He simply picked powerful models off the shelf, figured out how to talk to them in the right way, and turned them into weapons." Related:Cybersecurity Keeps Events 'Uneventful' Trim Details Six AI Jailbreaking Techniques Cato CTRL discovered the operation when a new account appeared on a Russian-language cybercrime forum on March 31 under the handle "Trim." The actor introduced himself and posted "a very detailed guide to AI jailbreaking" that laid out six named techniques for bypassing Anthropic's Claude Opus safety filters. They focused on manipulating how the models interpreted a user's intent and context, with the goal of persuading the AI to treat malicious requests as legitimate or harmless, according to the report. LOADING... The first, "Context Warming," starts with legitimate-looking security or coding requests to build trust before introducing malicious prompts. The second, "Black Box Principle," instructs the AI to analyze only code structure rather than intent, attempting to bypass its safety checks. "Ghost Reset," the third method Trim described, introduces a new chat after a refusal and rephrases the request, claiming the previous session was interrupted, while "Model Cascading" switches to alternative AI models if one refuses to comply, taking advantage of differences in safety guardrails. "Local Uncensored Models" uses self-hosted or locally run AI models with few or no safety restrictions when commercial models block requests, while the last jailbreak method, "Gray-Market API Access," obtains low-cost API keys from underground resellers to access commercial AI models without going through official channels. Monetizing the Methods At first, Trim didn't create a business out of the jailbreaking; however, nearly three months after publishing jailbreak techniques, the actor returned to the forum to launch "AI Pentest Checker," an automated Web vulnerability-scanning platform that integrates multiple AI models with a suite of offensive security tools to automate reconnaissance, vulnerability validation, exploitation reporting, and PDF-report generation. Related:Forgotten Bootloaders Expose Secure Boot Blind Spot The platform relies on a modified system prompt allegedly derived from a leaked Claude Fable 5 configuration to improve AI-assisted vulnerability escalation, which is a "significant escalation" of offensive tooling, according to the report. "As Mythos-class capabilities proliferate, whether through direct API access, key resellers, or leaked configurations, the offensive tooling built on top of them will grow in sophistication and accessibility," according to the report. Jailbreaks Become Attacks Indeed, the research shows how jailbreaking techniques, which have been around for some time, are not just the end of the story when it comes to dismantling AI model guardrails, Etay Maor, Cato Networks' vice president of threat intelligence, tells Dark Reading. They have now become the means for other malicious activities where AI is concerned. Related:Nigeria Deepens Cybersecurity Efforts as Cybercriminals See More Profits "This offering by Trim highlights the three main concerns when it comes to AI usage by threat actors," he says. Those concerns are: "The speed at which they can analyze, develop, and deploy; the scale at which they can do it in; and who now has access to tools and capabilities that once were only the hands of highly sophisticated groups or nation states," Maor says, adding that the last issue "is personally the most concerning aspect for me." The speed at which Trim turned his offensive knowledge into a repeatable service also stands out, Rickard Carlsson, CEO of application security provider Detectify, tells Dark Reading. This overall demonstrates how "AI changes the economics of cyber by reducing the time and expertise needed to move from experimentation to operation," he says. "The bigger story is that it allows existing techniques to scale faster and become accessible to a much wider group of actors," Carlsson says. How Should Defenders Respond Indeed, the rapid progression from sharing jailbreak methods to commercializing an AI-assisted offensive platform underscores how quickly threat actors can capitalize on generative AI. For this, defenders also need to change up their game to respond, even though "the fundamentals remain the same," Carlsson says. "Attackers still need an exposed asset, a weakness, and a path to something valuable," he says. However, what Trim's operation changes is "the speed and volume at which those paths can be discovered and revisited," which means visibility will become essential to security, Carlsson notes. "Security teams can no longer assume that an abandoned service, forgotten subdomain, or temporary exposure will remain unnoticed until the next security review," he says. "As AI accelerates discovery, organizations need a continuously updated understanding of what is Internet-facing and whether those assets represent real, exploitable risk." The findings also change detection, as Cato's Maor recommends that organizations build detection around behavioral sequences and early-stage reconnaissance, not individual suspicious requests. "In the past, the scanning was noisy, but now the individual requests may look less alarming," he explains. "However, the sequences for the attack are still there, so visibility becomes key to detecting these attacks, as the triaging across multiple steps is the key." About the Author Elizabeth Montalbano Contributing Writer Elizabeth Montalbano is freelance writer, editor, and journalist with 30 years of professional experience and a master's degree from Arizona State University. Her areas of expertise include enterprise technology, cybersecurity, business, and culture. During her long career, Elizabeth has lived and worked as a full-time journalist in Phoenix, San Francisco, and New York City. She specializes in news coverage and analysis, using her years of experience to look at the current state of cybersecurity with a critical gaze. She currently resides in a village on the southwest coast of Portugal, where in her free time she enjoys surfing, hiking with her dogs, growing plants, and playing and performing as a singer and musician. Want more Dark Reading stories in your Google search results? ADD US NOW More Insights Industry Reports The State of Cloud Security: The Latest Challenges How Organizations Are Managing Incident Response How Enterprises Are Developing Secure Applications Inside RSAC 2026: security leaders reveal the risks redefining your defense strategy Essential News & Insights from Black Hat USA 2025 Access More Research Webinars 0-Day to 10x Discovery: Security at the Speed of Mythos When AI Becomes an Insider: Rethinking Risk in Critical Infrastructure Governing the Agent; Identity Security in the Age of Autonomous AI Securing the AI Era: Shadow AI, AI Agents, and Why AI Detection and Response Changes Everything Practical Zero Trust Implementation on a Budget in the Age of Mythos More Webinars You May Also Like CYBER RISK Claude Mythos Fears Startle Japan's Financial Services Sector by Nate Nelson APR 30, 2026 CYBER RISK How Can CISOs Respond to Ransomware Getting More Violent? by James Doggett JAN 28, 2026 CYBER RISK US Cyber Pros Plead Guilty Over BlackCat Ransomware Activity by Alexander Culafi JAN 05, 2026 CYBER RISK Microsoft Exchange 'Under Imminent Threat,' Act Now by Arielle Waldman NOV 12, 2025 Editor's Choice VULNERABILITIES & THREATS Records Are Made to Be Broken: Patch Tuesday Raises Triage Stakes byJai Vijayan JUL 14, 2026 5 MIN READ PERIMETER 6 GHz Wi-Fi Flaws Could Disrupt Critical Systems byAlexander Culafi JUL 14, 2026 4 MIN READ CYBERSECURITY OPERATIONS 'Yellow Teams' Are Defining the Future of AI Security byNate Nelson JUL 13, 2026 6 MIN READ Want more Dark Reading stories in your Google search results? Keep up with the latest cybersecurity threats, newly discovered vulnerabilities, data breach information, and emerging trends. Delivered daily or weekly right to your email inbox. SUBSCRIBE LOADING... AUG 1-6 | MANDALAY BAY, LAS VEGAS USE CODE: DARKREADING & SAVE $200 ON A BRIEFINGS PASS OR $100 ON A BUSINESS PASS The premier cybersecurity event returns. GET YOUR PASS Discover More Black Hat Omdia Working With Us About Us Meet the Editors Advertise Reprints Join Us NEWSLETTER SIGN-UP Follow Us Copyright © 2026 TechTarget, Inc. d/b/a Infor

Share this article