Buried in the Opus 5 system card, not the headline eval scores: it's Anthropic's least prompt injectable model to date.


The receipts:
> tested across prompt-injection evals, not just one scenario
> held up against active red team attempts
> flagged by Anthropic's own lead as the result he's most excited about, over the benchmarks
Prompt injection is the attack that turns your agent against you. Every tool call on untrusted data is exposure. Moving that number is worth more than a point on a coding bench.
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned