
Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents
Quick Answer
Anthropic's Opus 5 demonstrates near immunity to prompt injection attacks, achieving a 0% success rate in browser scenarios across 129 tests.
Quick Take
In general tests, it reduced the success rate from 5.5% (Opus 4.8) to 2.0%, outperforming competitors like Mythos 5 and Fable 5.
Key Points
- Opus 5 achieves a 0% success rate in prompt injection for browser agents.
- Auto Mode combines two defense layers to enhance security against attacks.
- Without Auto Mode, Opus 5's success rate is 3.7%, while Sonnet 5 is at 0.93%.
- Gray Swan's benchmark shows Opus 5 leads with a 2.0% success rate after 15 attempts.
- OpenAI previously acknowledged that prompt injection may never be fully resolved.
Source Excerpt
Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3. 7 percent. If these numbers hold up in practice, Anthropic may have cracked one of the biggest security problems facing AI agents that operate in browsers.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

