Claude Opus 5 Tops Several Benchmarks While Undercutting Fable 5 on Cost
Key Takeaways
- •Claude Opus 5 scored 61 on the Artificial Analysis Intelligence Index, narrowly ahead of Claude Fable 5, GPT-5.6 Sol, Kimi K3, and Claude Opus 4.8.
- •In coding tests, Opus 5 shared first place on the Artificial Analysis Coding Index and matched GPT-5.6 Sol with an 89 percent score on Terminal-Bench v2.1 at the max tier.
- •Factual accuracy remains a weaker area, with Opus 5 trailing Fable 5 on AA-Omniscience and its hallucination rate rising by 14 points to 50 percent.
- •Epoch AI rated Opus 5 slightly below Fable 5 overall, while GPT-5.6 Sol led both the overall Epoch Capability Index and the software engineering category.
- •Vals.ai found that Opus 5’s high reasoning tier performed best on Vibe Code Bench, outperforming more expensive xhigh and max tiers.

Anthropic's Claude Opus 5 is the most capable AI model currently available according to several benchmark results, with particularly strong performance in analytical quality, coding, and knowledge-work tasks. The model matches or exceeds Claude Fable 5 across many tests while generally costing less, though evaluations also show a higher hallucination rate when the model answers despite uncertainty.
On the Artificial Analysis Intelligence Index, Claude Opus 5 scored 61. The index combines nine tests spanning knowledge work, coding, scientific reasoning, and factual accuracy. That result places Opus 5 narrowly ahead of Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, and Claude Opus 4.8 at 56. Artificial Analysis worked with Anthropic to test the model before its public release.
In coding benchmarks, Claude Opus 5 at the "xhigh" reasoning level, used with Claude Code, shares first place on the Artificial Analysis Coding Index. The index measures how well AI models can independently complete programming tasks, including identifying and fixing bugs. On Terminal-Bench v2.1, which evaluates agents acting as autonomous engineers in real terminal environments, Opus 5 scored 89 percent at the "max" tier, matching the previous leader, GPT-5.6 Sol.
For scientific reasoning, Opus 5 scored 53 percent on Humanity's Last Exam, a difficult multi-domain academic knowledge test, tying Claude Fable 5. On CritPt, a physics benchmark developed by researchers at Argonne National Laboratory and UIUC, Opus 5 again matched Fable 5 but remained behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra.
Factual accuracy is still a weaker area. On AA-Omniscience, which tests the accuracy of a model's knowledge claims, Opus 5 improved by 7 points over Opus 4.8 but continued to trail Fable 5. The model also responds more often when it is uncertain, raising its hallucination rate by 14 points to 50 percent. That trade-off matters for deployments where users want models to complete long, complex tasks but still need reliable refusal or uncertainty handling when the answer is not supported.
Epoch AI results show a close frontier-model race
Epoch AI has also evaluated Claude Opus 5. The research institute gave the model an overall Epoch Capability Index score of 159, slightly below Fable 5 at 161. In the software engineering subset, known as SWE-ECI, Opus 5 tied Fable 5 at 161 and outperformed GPT-5.6 Terra and Claude Opus 4.8. GPT-5.6 Sol leads both the overall index and the software engineering category.
The results point to a narrow performance gap among frontier models. No single model shows a clear lead across all categories. The finding is consistent with the argument that AI models may become increasingly commoditized, a view also discussed in connection with comments by Microsoft CEO Satya Nadella. It also shows why buyers and developers increasingly compare models by workload, latency, reliability, and cost rather than by a single headline score.
Lower reasoning tiers can offer better value in coding
The average Intelligence Index task costs $2.03 with Opus 5. That is lower than Claude Fable 5 with fallback at $2.75, but higher than Opus 4.8 at $1.80 and Sonnet 5 at $1.53. At the "high" and "xhigh" reasoning tiers, however, Opus 5 outperforms both Opus 4.8 and Sonnet 5 while costing less.
Vals.ai tested Claude Opus 5 across all five reasoning tiers using Vibe Code Bench, a programming-task benchmark. Scores rose from 76.7 percent at "low" to 82 percent at "medium" and 89.8 percent at "high." Performance then dipped at the two highest tiers: "xhigh" scored 88.3 percent and "max" scored 88.4 percent, despite substantially higher costs.
Vals.ai found that the highest tiers tend to generate more complex solutions, which more frequently contain errors. The "high" tier produced simpler solutions that met requirements more reliably.
Terminal-Bench 2.1 shows a similar pattern. The "high" tier outperformed "max" because the model spends more time on each attempt at the highest tier, leaving fewer attempts within the benchmark's time limit. That aligns with Anthropic's guidance, which sets "high" as the default tier in both the API and Claude Code.
Token pricing remains $5 per million input tokens and $25 per million output tokens. Cache writes cost $6.25 per million tokens with a five-minute lifetime, while cache hits cost $0.50 per million tokens. Those prices make output length, retry behavior, and prompt caching important parts of the effective cost for agentic coding and document-heavy workflows.
Opus 5 leads in knowledge-work evaluations
Opus 5 performs especially strongly on the AA-Briefcase benchmark, which measures how well AI models handle common office tasks such as writing research reports, creating presentations, and analyzing spreadsheets based on thousands of input files. The benchmark scores performance across correctness, analytical quality, and presentation quality, then converts those results into an Elo rating similar to chess rankings.
At max reasoning, Opus 5 reached an Elo score of 1720, which is 146 points ahead of Claude Fable 5 at 1574. Its three highest tiers — max, xhigh, and high — took the top three positions. Together with Fable 5, Sonnet 5, and Opus 4.8, Anthropic models now account for the large majority of the top 10 positions.
Cost per task fell 20 percent to $17.79, compared with $22.30 for Fable 5. The "xhigh" version costs $14.26 per task, while "high" costs $10.41, less than half the cost of Fable 5. Both tiers still ranked ahead of Fable 5 by Elo. At medium performance, Opus 5 posted an Elo score of 1470, just behind GPT-5.6 Sol at max with 1505. At low performance, Opus 5 scored 1223 Elo, slightly below GLM-5.2 at max with 1254.
The largest gains for Opus 5 appear in analytical quality. At "max," the model reached an Analytical Quality Elo of 2016, nearly 300 points ahead of Fable 5. Its Rubric Pass Rate, which measures how often the model satisfies predefined quality standards, was 58 percent at "max," 57.2 percent at "xhigh," and 56 percent at "high."
Presentation quality was less dominant. Opus 5 recorded a Presentation Elo of 1628, about 40 points behind GPT-5.6 Sol at "max," which scored 1666.
Higher performance also requires more time. At "max," Opus 5 takes more than 36 minutes per task and averages 103 passes, roughly 50 percent longer than Opus 4.8, which takes 24 minutes and averages 55 passes. For teams evaluating the model, the benchmark results suggest that the best tier can depend on whether the task rewards maximum analysis, faster turnaround, or lower per-task cost.