
Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents
Quick Answer
Anthropic's Opus 5 demonstrates near immunity to prompt injection attacks, achieving a 0% success rate in browser scenarios across 129 tests.
Quick Take
In general tests, it reduced the success rate from 5.5% (Opus 4.8) to 2.0%, outperforming competitors like Mythos 5 and Fable 5.
Key Points
- Opus 5 achieves a 0% success rate in prompt injection for browser agents.
- Auto Mode combines two defense layers to enhance security against attacks.
- Without Auto Mode, Opus 5's success rate is 3.7%, while Sonnet 5 is at 0.93%.
- Gray Swan's benchmark shows Opus 5 leads with a 2.0% success rate after 15 attempts.
- OpenAI previously acknowledged that prompt injection may never be fully resolved.
📖 Reader Mode
~2 min readAnthropic says Opus 5 is nearly immune to prompt injections in its own software. Prompt injection, where an attacker slips past an AI model's instructions through manipulated inputs like hidden text on a webpage, fails against Opus 5 in almost every case. For browser agents, the attack success rate hit zero percent across 129 test scenarios, per the system card. That's a big deal given that OpenAI admitted in December that prompt injection may never be fully solved. In a general prompt injection test by security firm Gray Swan, the success rate after 15 attempts dropped from 5.5 percent (Opus 4.8) to 2.0 percent.

That zero percent rate only holds with Auto Mode turned on in products like Claude Cowork. Auto Mode stacks two defense layers. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before execution. An attacker has to beat both independently. Without them, Opus 5 sits at 3.7 percent, and Sonnet 5 actually does better at 0.93 percent. Only the combination of model and protective software pushes the rate to zero.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

