Opus 5 topped the intelligent ranking by a narrow margin, with a 26% lower single-task cost than Fable 5

robot
Abstract generation in progress
According to Beating monitoring by Dongcha, third-party evaluation firm Artificial Analysis has released Claude Opus 5’s results. It scored 61 points on the smart index covering 9 tests, narrowly ahead of Fable 5’s 60 points. GPT-5.6 Sol scored 59 points, and Kimi K3 scored 57 points.

Opus 5’s average cost per task is $2.03, which is 26% lower than Fable 5’s $2.75. It topped both the GDPval-AA v2 and AA-Briefcase knowledge-work evaluations, and, together with Claude Code, it also jointly took first place in the programming agent index. Terminal-Bench v2.1 scored 89%, roughly matching GPT-5.6 Sol.

The model offers 5 tiers of reasoning strength. From low to max, the output Token count differs by about 8 times, and the GDPval-AA v2 results differ by 407 Elo. Users can trade more Tokens for stronger performance, or proactively reduce costs.

The weaknesses are also clear. Opus 5’s factual knowledge still lags behind Fable 5. In the AA-Omniscience test, its hallucination rate rose to 50%, which is 14 percentage points higher than Opus 4.8. The cost-effectiveness in the lower reasoning tiers still trails the GPT-5.6 series slightly.
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned