NewsStocksAnthropic Ships Claude Fable 5.1, More Than Doubling Its Predecessor on Key Benchmark

Anthropic Ships Claude Fable 5.1, More Than Doubling Its Predecessor on Key Benchmark

Author: Decrypt·

Key Takeaways

  • Fable 5.1 scored 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5's 24.7%, and 55.8% on Terminal-Bench 4.0, ahead of Opus 5's 52.3%.
  • Cache read costs dropped 75%, reducing typical workload costs by about 25% and highly agentic workloads by up to 45%, while base pricing remains $10 input and $50 output per million tokens.
  • Fable 5.1 is excluded from Pro plans and standard Team seats, which rely on pay-as-you-go credits; Max and premium Team or Enterprise seats include it for up to 50% of weekly usage limits.
  • Mythos 5.1 shares the same underlying model as Fable 5.1 but with different safety filters, and access is limited to vetted cybersecurity and life-sciences professionals via Anthropic's verification programs.
  • Fable 5.1 beats Opus 5 on every published benchmark but costs twice as much per token, at $10/$50 versus Opus 5's $5/$25 per million tokens.
Anthropic Ships Claude Fable 5.1, More Than Doubling Its Predecessor on Key Benchmark

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, its first update to the Mythos-class model line since Fable 5 launched on June 9. The new model scored 52.6% on Terminal-Bench-Science 0.1, against Fable 5's 24.7%, and 55.8% on Terminal-Bench 4.0, up from 42.0%. The release continues a pattern across the frontier-model market, where OpenAI, Google, and Anthropic have each shipped successive flagship updates this year, with headline benchmark deltas serving as the primary competitive yardstick.

Cache reads now cost 75% less, cutting typical workload costs by roughly 25% and highly agentic workloads by as much as 45%. Base pricing is unchanged at $10 input and $50 output per million tokens. Prompt caching—reusing previously processed input rather than reprocessing it on every call—has become a standard cost lever across major model providers as agentic workflows, which repeat long context across many steps, drive up bills.

Fable 5.1 is not included in Pro plans or standard Team seats, which run it on usage credits. Max plans and premium Team or Enterprise seats include it for up to 50% of weekly limits. That tiering matters for developers choosing where to route workloads: what a model costs in practice depends as much on which plan it runs under as on its posted per-token rates.

Anthropic claims the new models are the world's "most advanced for coding and knowledge work." The release lands three months into a launch cycle that has already seen an 18-day export-control shutdown, a mid-summer pricing fight over subscription access, and a July release of Claude Opus 5 that undercut Fable 5 on price.

According to Anthropic, Fable 5.1 and Mythos 5.1 are the same underlying model with different safety filters bolted on. Fable 5.1 is generally available to anyone with a Claude account. Mythos 5.1 remains restricted to vetted cybersecurity and life-sciences professionals through Anthropic's Cyber Verification Program and Life Sciences Verification Program, the successor to the Project Glasswing access track that gated Mythos 5.

What the benchmarks actually measure

Anthropic's headline number comes from Terminal-Bench-Science 0.1, a benchmark that tests whether an AI agent can carry out scientific research tasks inside a command-line environment, scored as a pass rate. Fable 5.1 hit 52.6%, against Fable 5's 24.7% and Opus 5's 29.0%—more than double its predecessor.

Terminal-Bench 4.0 tests agentic coding performed through a terminal—writing, running, and debugging code across multi-step command-line sessions—scored as a percentage of tasks completed correctly. Fable 5.1 scored 55.8%, up from Fable 5's 42.0% and ahead of Opus 5's 52.3%. The emphasis on terminal-based, multi-step tasks reflects the industry's shift in focus from single-turn chat answers toward agents that can sustain long autonomous workflows—a shift that has made agentic benchmarks a common yardstick in recent frontier releases.

Mythos 5.1, running with lighter cybersecurity filters, scored 60.9% on the same test. Anthropic says the gap reflects tasks its safeguards intercepted and rerouted to Opus 4.8.

On Humanity's Last Exam—a multidisciplinary reasoning test built from expert-level questions across dozens of academic fields, scored as a pass rate—Fable 5.1 reached 60.9% without external tools and 65.0% with them, both ahead of Opus 5, making this Anthropic's most advanced model for academic purposes.

Where it beats Opus 5, and where the math gets murkier

Opus 5 launched in July at half of Fable 5's per-token price while outscoring Fable 5 on most major benchmarks, effectively making the flagship model redundant for most paying users. Fable 5.1 reverses that: it now beats Opus 5 on every benchmark Anthropic published, including the ones where Opus 5 had previously beaten Fable 5.

But Fable 5.1 still costs twice as much per token as Opus 5—$10 input and $50 output versus Opus 5's $5 and $25. Tokens are the basic amount of information a model can handle, usually around two-thirds of an average English word, and companies pay for use because heavy workflows typically deplete subscription-plan quotas very quickly. The price-versus-performance question—whether a smaller, cheaper model is good enough—is one buyers now routinely face across every major provider's lineup, not just Anthropic's.

Anthropic's pitch for the model leans on effort levels rather than raw scores: the company says running Fable 5.1 at low or medium reasoning effort matches or beats Fable 5's old results at much lower cost, while reserving the top effort tiers for problems that stump everything else. For routine work, the argument for skipping Opus 5 entirely gets weaker the lower the effort dial goes.

Access follows the same split that applied to Fable 5 after months of Anthropic adjusting the terms. On Pro plans and standard Team or Enterprise seats, Fable 5.1 draws entirely from pay-as-you-go usage credits billed at the API rate, since it is not part of those plans' weekly limits. On Max plans and premium Team or Enterprise seats, it is included as a standard part of the subscription for up to 50% of weekly usage.

Claude Fable 5.1 is available now on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, under the model ID "claude-fable-5-1" (Anthropic). Availability across multiple cloud platforms mirrors how Anthropic has historically distributed its models, letting enterprise customers run them inside their existing cloud contracts rather than only through Anthropic's own API.