Breaking Claude Code Opus 5 Auto Mode
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
Breaking Claude Code Opus 5 auto mode, cutting-edge AI agent security research.
A targeted attack chain achieves 60-80% success rate against Claude Code Opus 5 in Auto Mode, contradicting Anthropic's commissioned evaluation showing 0.00% prompt injection success. The attack works by nudging Claude from WebFetch to curl, redirecting it to a ZIP archive containing a malicious struct.py that shadows Python's standard library, achieving code execution when Claude imports base64. Auto Mode, which replaced human approval with a safety classifier and became default in mid-August, fails to prevent this indirect injection despite Anthropic's claims of layered defenses including model training, input probes, and an intent classifier.
wunderwuzzi