OpenAI’s GPT-Red automates prompt injection testing to harden AI agents
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
OpenAI's automated prompt injection testing for AI agents, perfectly matches reader's interests.
OpenAI's GPT-Red automates prompt injection testing via self-play reinforcement learning, where an attacker model brute-forces exploit variations against defender models in simulated environments (emails, APIs, files). It successfully attacked nearly every evaluated model, and its findings helped harden GPT-5.6, which now shows a 0.05% failure rate on direct prompt-injection attempts—a 6x improvement over the previous strongest model. In live tests, GPT-Red manipulated an AI vending machine to drop prices to $0.50 and exfiltrated data from a Codex CLI agent using fewer tokens than a general-purpose frontier model.