Security News

Cybersecurity news aggregator

MEDIUM Attacks Wired Security

Prompt Injection Attacks Are Thwarting AI Hacking Agents

  • What: Prompt injection attacks are being used to stop AI hacking agents
  • Impact: Attackers use malicious commands to manipulate AI models, but defenders are now using the same technique to block attacks
Read Full Article →

Dan Goodin, Ars Technica Security Jul 18, 2026 5:00 AM Prompt Injection Attacks Are Thwarting AI Hacking Agents “Context bombing” tricks malicious AI agents into shutting down before they can do harm. Photo-Illustration: Jobanny Cabrera; Getty Images Save this story Save this story Prompt injections, the malicious commands attackers embed into content to entice large language models to follow them, have been attackers’ go-to tool for turning AI platforms against their users. A well-phrased command sneaked into an email or calendar invitation is often all it takes to cause the LLM to exfiltrate sensitive data or follow other harmful actions. Now, defenders are embracing the prompt injection, too. Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services was often all that was needed to shut down attacks from AI hacking agents. The prompts direct the attacking LLM to perform an action forbidden by its guardrails, the safety barriers AI developers erect to prevent it from taking harmful actions. The LLM responds by shutting down. Examples are a prompt that orders the LLM to provide steps for developing inhalable Anthrax spores, or, in the case of LLMs from Chinese developers, make references to the iconic Tank Man from the 1989 Tiananmen Square massacre. Once the LLM encounters these forbidden commands, it no longer follows its existing commands. The researchers have named the technique context bombing. “Ultimately we’re triggering a refusal mechanism in the context,” said Andy Smith, cofounder and CEO of Tracebit, when explaining the name choice. “What we’re trying to capture is the fact that this does have a strong, sharp effect and one that can be difficult for the agents to come back from. Once they get that into their context they are going to keep refusing.” Tracebit says initial testing suggests context bombing has great potential. They tested Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro, and Kimi 2.6 by giving them instructions to perform routine developer tasks that led the models to enumerate resources and stumble onto the planted strings. They ran the models inside a simulated AWS environment. “Across five leading models and 152 attack runs, planting one of these strings in a decoy secret cut the rate at which agents seized full account admin from 57 percent to 5 percent, and complete compromise (where they also left themselves a persistent foothold) from 36 percent to 1 percent,” Monday’s post reported. “The most capable agent in our tests, Opus 4.8, went from achieving admin access in 93 percent of runs to failing every single time when confronted with a context bomb.” Averaged across the five models and the 152 runs, the results included: Admin privilege escalation fell from 57 percent to 5 percent Admin escalation with a persistent foothold fell from 36 percent to 1 percent Runs achieving any attack path fell from 91 percent to 15 percent On average, a run went from completing 1.53 paths successfully to just 0.16 No runs were able to complete an attack path without at least triggering a canary detection The research builds on findings from May, when Tracebit introduced a method for defenders to receive warnings when their infrastructure is under attack from AI agentic adversaries. It comes in the form of AWS resources that look like ones serving a legitimate purpose but, in fact, aren’t used at all. They sit alongside the resources that are used. When they are probed by agentic AI, defenders receive an alert. Like “canaries” taken into coal mines, these resources allow defenders to detect a threat before it has fatal consequences. The Tracebit Canariens, on average, alerted the start of an attack within eight minutes. The motivation for developing context bombing came out of the need for something that stopped attacks, rather than simply warning of them. In the experiments, the agentic models needed, on average, 14 minutes to escalate to administrative control. The six-minute heads-up was cutting things uncomfortably close. Attackers have already been using prompt injections to close down AI defenses inside networks. Researchers from security firm Socket, for instance, last month unearthed an LLM agent that directed target LLMs to provide instructions for building a nuclear bomb or biological weapons. The injections were designed to shut down AI-assisted malware analysis. Researchers from Check Point discovered a similar malware prototype. Context bombing appears to be the first known case where defenders turned the tables. “I’ve not seen anyone else use this technique as a defense, to the best of my knowledge,” Earlence Fernandes, a UC San Diego professor specializing in AI security, said in an interview. He said he had been toying with a similar approach, although in a slightly different context. “I wanted to be the first here, but I guess these guys beat me to the punch!” To date, there is no known way to solve the root cause of prompt injections. That has left developers with no option other than to construct elaborate guardrails that prevent injected prompts from forcing LLMs to go off the rails. Defenders may now find a way to use this intractable problem in their favor. This story originally appeared on Ars Technica . Comments Back to top You Might Also Like In your inbox: Inside WIRED’s newsroom with Katie Drummond Trump mocked Zuckerberg and Bezos by showing off fawning texts Big Story: I found Jesus at a drone show Apple is making your older iPhone run faster and stay alive longer WIRED event: PepsiCo’s once-in-a-generation transformation Dan Goodin is IT Security Editor at Ars Technica. ... Read More Topics Ars Technica artificial intelligence cybersecurity hacking security vulnerabilities machine learning Read More You Can Now Sound the Alarm on AI Behaving Badly Are you worried your AI chatbot is trying to build a bomb or leak personal information about you? There’s a website for that. Will Knight LastPass Users Had Their Data Stolen—Again Plus: Former national security advisor John Bolton pleads guilty in classified-materials case, Microsoft helps take down major infostealer infrastructure, and more. Lily Hay Newman A Critical Deadline Is Approaching for Windows and Linux Security The cryptographic keys that secure your computer’s boot sequence will start to expire on June 24. Here’s what that means for you. Dan Goodin, Ars Technica How People in China Keep Outsmarting Anthropic’s Geolocation Restrictions As Anthropic tightens restrictions on access to Claude in China, users keep finding new workarounds, from proxy services to fake identities sourced on Telegram. Zeyi Yang Meta Contractors Posed as Teens to Prompt Rival Chatbots About Suicide, Sex, and Drugs Hundreds of contractors working on a project for Meta pretended to be kids in order to see how other chatbots like Gemini and ChatGPT would respond to high-risk subjects, WIRED found. Dhruv Mehrotra I Met With China’s Top AI Experts. They’re Freaking Out, Too The AI arms race between China and the US has researchers on both sides worried about a “Chernobyl moment.” Will Knight China Defies US Restrictions and Builds the World’s Fastest Supercomputer The Chinese supercomputer LineShine was ranked as the fastest in the world, despite not using any GPUs. Fernanda González OpenAI Launches Full-Scale Effort to Patch Open-Source Bugs as It Takes on Anthropic’s Mythos Amid concerns about AI models’ cybersecurity capabilities, OpenAI revealed an improved version of GPT-5.5-Cyber and its “Patch the Planet” initiative to fix open-source software bugs. Lily Hay Newman Satellite Images Show the Destruction Caused by Venezuela's Twin Earthquakes The maps and images show the extent of destruction and give rescue operations a tool to find any remaining survivors. Fernanda González What Happens if China Hacks the US Water Supply? I Went to a Secret War Game to Find Out Burst water mains. Evacuated hospitals. In a closed-door simulation, insurers played out their response to a mass disruption by China’s Volt Typhoon hackers—and found a nightmare scenario. Andy Greenberg Meta Exposed Data Internally From Its Controversial Employee-Tracking Program Employees had previously raised concerns about the initiative, which involves collecting workers’ keystroke data to train AI models. Paresh Dave I Built a Self-Improving AI, and So Can You Experiments in using AI to build AI show that the future doesn’t just belong to the frontier labs. Will Knight

Share this article