Opus 5 narrowly tops the intelligence chart, with 26% lower cost per single task than Fable 5

robot
Abstract generation in progress
According to Beating monitoring by Beating Insight, the third-party evaluation firm Artificial Analysis released Claude Opus 5’s results. It scored 61 points on the intelligent index across 9 tests, narrowly ahead of Fable 5’s 60. GPT-5.6 Sol scored 59, and Kimi K3 scored 57.

Opus 5’s average cost per task is $2.03, which is 26% lower than Fable 5’s $2.75. It topped the knowledge work evaluations in GDPval-AA v2 and AA-Briefcase, and, together with Claude Code, it also tied for first place in the programming agent index. Terminal-Bench v2.1 scored 89%, roughly matching GPT-5.6 Sol.

The model offers 5 tiers of inference strength. From low to max, the output Token difference is about 8 times, and the GDPval-AA v2 score difference is 407 Elo. Users can trade more Tokens for stronger performance, or proactively lower costs.

The shortcomings are also evident. Opus 5’s factual knowledge still lags behind Fable 5. In the AA-Omniscience test, its hallucination rate rose to 50%, 14 percentage points higher than Opus 4.8. The value-for-money at low inference tiers also remains slightly inferior to the GPT-5.6 series.
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned