Anthropic Says Claude Opus 5 Nears Fable 5 Performance at Half the Token Price
Key Takeaways
- •Claude Opus 5 retains the same token pricing as its predecessor at $5 per million input tokens and $25 per million output tokens while becoming the default model on Claude Max and the most capable option on Claude Pro.
- •Opus 5 leads on several benchmarks including Frontier-Bench v0.1 agentic terminal coding at 43.3% and ARC-AGI-3 at 30.2%, but trails GPT-5.6 Sol on DeepSWE and lags Mythos 5 on cybersecurity tasks.
- •The model can self-correct through iterative steps and write code to build its own tools, demonstrated when it independently created a computer vision pipeline to reconstruct a 3D machine part from a drawing.
- •Opus 5's cyber safety classifiers trigger roughly 85% less often than those used with Fable 5, responding to developer criticism that Fable 5's frequent interventions disrupted legitimate coding and security research.
- •The max effort setting produces slightly worse results than the second-highest setting on two benchmarks despite costing more, indicating that higher computational effort does not guarantee improved model outputs.

Anthropic is positioning its new flagship model, Claude Opus 5, as a lower-cost alternative to Claude Fable 5, saying the model approaches Fable 5's performance while beating it on several benchmarks. The release comes as AI labs compete intensively on both capability and cost, with enterprise customers increasingly weighing price-performance tradeoffs when selecting models for production workloads.
The company says Opus 5 leads tests in agentic coding and knowledge work—two of the areas where businesses are deploying AI most heavily. On ARC-AGI-3, a benchmark designed to measure novel problem-solving, Opus 5 scores 30.2 percent, nearly four times the result posted by GPT-5.6 Sol.
Anthropic also says Opus 5 can check and improve its own work through iterative steps and can write code to build tools when it needs them, capabilities that matter for autonomous agent workflows where models must complete multi-step tasks with minimal human intervention.
Opus 5 keeps Opus 4.8 pricing
Anthropic is introducing Opus 5 as pricing pressure increases from GPT-5.6 Sol and Chinese competitors. The new model is intended to narrow the price-performance gap with the more expensive Fable 5. Opus 5 becomes the default model on Claude Max and the most capable model offered on Claude Pro.
The model keeps the 1 million-token context window and the same token rates as its predecessor, Opus 4.8. Anthropic charges $5 per million input tokens and $25 per million output tokens for Opus 5. A new Fast Mode raises speed by 2.5x, but it doubles the price.
Base token rates do not fully reflect total task costs, because token efficiency can vary between models. Opus 4.7 cost 30 to 40 percent more per task than Opus 4.6, even though both models used the same base rates. A similar pattern recently appeared with Claude Sonnet 5. This distinction matters for organizations running large-scale deployments, where per-task cost differences compound across millions of API calls.
Higher effort settings may cost more without improving results
Users can balance performance and token usage through five effort settings: low, medium, high, xhigh, and max. Anthropic says Opus 5 provides better value than its predecessor at every effort level.
In its prompting guide, Anthropic recommends broad use of the "low" and "medium" settings. The company says those settings deliver good results with a fraction of the token use and latency while outperforming the same settings on earlier Opus models. For coding and agentic tasks, Anthropic still recommends starting with "xhigh."
Opus 5 performs slightly worse at the max effort setting than at the second-highest setting on two benchmarks, even though max costs more. The decline appears on Frontier-Bench v0.1 and the Artificial Analysis Coding Agent Index. The finding underscores that higher computational effort does not guarantee better outputs, a consideration for teams optimizing cost versus quality.
Anthropic benchmarks show Opus 5 ahead in agentic coding
According to Anthropic's own benchmark results, Opus 5 sets records across several evaluations. On Frontier-Bench v0.1, it reaches 43.3 percent on agentic terminal coding, ahead of Fable 5 at 33.7 percent, GPT-5.6 Sol at 34.4 percent, and Opus 4.8 at 21.1 percent.
On the GDPval-AA v2 knowledge work benchmark, Opus 5 leads with an Elo score of 1,861. Fable 5 scores 1,747, while GPT-5.6 Sol scores 1,736.
Opus 5 does not lead in every test. On DeepSWE v1.1, which measures agentic coding, GPT-5.6 Sol ranks first with 72.7 percent, followed by Fable 5 at 69.7 percent and Opus 5 at 68.8 percent. On health tasks and legal benchmarks, Fable 5 and Mythos 5 outperform Opus 5, respectively. The mixed results highlight how model leadership varies by domain, making model selection task-dependent for enterprise buyers.
The ARC-AGI-3 result is one of the largest outliers in Anthropic's benchmark set. ARC-AGI-3 evaluates novel problem-solving without relying on memorized patterns. Opus 5 scores 30.2 percent, compared with 1.5 percent for Opus 4.8 and 7.8 percent for GPT-5.6 Sol. That places Opus 5 nearly four times ahead of the next-best model in the cited results. Anthropic does not list a Fable 5 result for this benchmark, and it remains unclear whether such a large benchmark lead will translate into real-world use.
Opus 5 trails Mythos 5 on cybersecurity tasks. Anthropic says it deliberately did not train Opus 5 on cyber tasks, as was also true for its predecessor. The model comes close to Mythos 5 in finding vulnerabilities, but performs much worse when asked to exploit them.
Anthropic also says Opus 5 has improved at generating visual outputs and analyzing visual content, including charts and diagrams.
Anthropic says Opus 5 can build its own tools
Anthropic describes Opus 5 as substantially better at reviewing its own work and improving it through iteration. This type of self-correction is a growing focus for AI labs building autonomous agents, where the ability to detect and fix errors without human guidance determines whether a model can handle complex, open-ended tasks.
In one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to create a 3D model in FreeCAD. The model was intentionally not given a direct way to view the drawing.
According to Anthropic, Opus 5 wrote its own computer vision pipeline to extract geometry from raw pixels and then reconstructed the complete machine part. The company says no other model solved the task after five attempts.
Opus 5 also worked on a real bug in a popular open-source package manager. Anthropic says the model identified the root cause and fixed an edge case that the community patch had missed. A competing model fixed only the surface-level symptom before marking the bug as resolved.
Anthropic also says an engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Earlier models were unable to complete the task, even when given detailed plans.
Cyber filters trigger less often than on Fable 5
Anthropic says Opus 5's safety setup permits source code vulnerability research while blocking binary-based vulnerability scanning, penetration testing, and exploit generation. Its cyber classifiers trigger about 85 percent less often than those used with Fable 5. Fable 5's frequent interventions had drawn heavy criticism from developers who reported workflow interruptions during legitimate coding and security research tasks.
Blocked requests in Claude.ai, Claude Code, and Claude Cowork fall back to Opus 4.8 by default, the same approach Anthropic used with Fable 5.
Anthropic calls Opus 5 the most capable generally available model for scientific research. The company says it improves on Opus 4.8 across all life sciences evaluations, with notable gains in organic chemistry and protein-related tasks. In organic chemistry, Opus 5 improves by 10.2 percentage points on deriving molecular structures from spectroscopy data. In protein-related tasks, it improves by 7.7 percentage points.
Anthropic is also releasing two beta features alongside Opus 5. Mid-Conversation Tool Changes on the Claude Platform let developers change available tools during a conversation without invalidating the prompt cache. Automatic Fallbacks on the API automatically route blocked requests to a different model. Both features reflect Anthropic's effort to address developer friction points around workflow continuity and reliability.