Prompt Injection Attacks Are Thwarting AI Hacking Agents
Tracebit's 'context bombing' plants forbidden prompts alongside AWS secrets to trigger refusal mechanisms in LLM hacking agents, cutting admin privilege escalation from 57% to 5% across five models including Opus 4.8 and Gemini 3.1 Pro. The technique, building on their canary detection approach, reduced complete compromise (persistent foothold) from 36% to 1% in 152 simulated attack runs. This defensive prompt injection exploits the same vulnerability attackers use, turning AI agents' guardrails against them.