Deep Tide TechFlow message: On July 25, Anthropic announced that its Opus 5 model is almost immune to prompt-injection attacks in browser agent scenarios. In 129 test scenarios, the attack success rate was zero. In the Gray Swan general prompt-injection test, after 15 attempts the success rate dropped from 5.5% with Opus 4.8 to 2.0%. A zero success rate was achieved only when products such as Claude Cowork enabled Auto Mode, a mode that overlays two layers of defense—input scanning and execution interception. Prompt injection is considered one of the biggest security risks facing AI agents, and this improvement may indicate that the issue has been effectively mitigated in specific scenarios.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned