Anthropic's Opus 5 Achieves Zero Success Rate Against Prompt Injection Across 129 Browser Test Scenarios

Anthropic announced on July 25 that its Opus 5 model achieved zero success rate against prompt injection attacks in 129 browser testing scenarios. In the Gray Swan universal prompt injection benchmark, the model's success rate dropped to 2.0% from Opus 4.8's 5.5% across 15 attempts. Zero success rates were achieved exclusively when Auto Mode was enabled in Claude Cowork and similar products, which combines input scanning and execution blocking defenses.
Disclaimer: The information on this page may come from third-party sources and is for reference only. It does not represent the views or opinions of Gate and does not constitute any financial, investment, or legal advice. Virtual asset trading involves high risk. Please do not rely solely on the information on this page when making decisions. For details, see the Disclaimer.
Comment
0/400
No comments