Anthropic Launches Claude Sonnet 5.5, Which Outcodes Opus 5.5 at Half the Price
Key Takeaways
- •Claude Sonnet 5.5 debuts at unchanged pricing of $2 per million input tokens and $10 per million output tokens, half the rates Anthropic charges for Opus 5.5.
- •The new model beat Opus 5.5 on Terminal-Bench 4.0 coding tests, scoring 70.6% versus 66.4% in Anthropic's results and 63.6% versus 59.6% in Artificial Analysis' independent run.
- •At maximum effort, Sonnet 5.5 generated roughly 193,000 tokens per task, the most Artificial Analysis has measured, costing $7.60 per task—about 50% more than Sonnet 5.
- •Anthropic says the model runs more than 30% faster than Sonnet 5 and uses nearly a third fewer tokens per task, making it cheaper to run despite identical list prices.
- •Claude Haiku 5.5, designed for high-volume, cost-sensitive applications, is expected in the coming weeks to round out Anthropic's 5.5-generation lineup.

Anthropic released Claude Sonnet 5.5 on Monday, the newest version of its middle-tier AI model and an upgrade to Sonnet 5, which debuted in June. Pricing is unchanged: $2 per million input tokens and $10 per million output tokens—half of what the company charges for Opus 5.5.
Despite the lower price, the model is already outscoring its larger sibling in coding. On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6% against 66.4% for Opus 5.5, according to Anthropic. Artificial Analysis, an independent testing firm, ran its own version of the benchmark and also placed Sonnet 5.5 ahead, 63.6% to 59.6%. The firm nonetheless ranks it second overall behind Opus 5.5 on its leaderboard and notes that it used more tokens per task than any model it tested.
Anthropic says the new model runs more than 30% faster than its predecessor. "Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It's also got a sharp eye for design," the company wrote.
Tokens are the chunks of text an AI model reads and writes—slightly shorter than a word—and AI companies bill by the million. While the sticker price matches Sonnet 5's, the new model uses nearly a third fewer tokens per task, which means it ends up being cheaper to run than its predecessor despite the identical rates.
Where the model truly shines is coding. Terminal-Bench 4.0 measures whether an AI agent can complete complex professional tasks by typing commands on its own, scored as the share of tasks completed. Sonnet 5.5 hit 70.6%. Opus 5.5 scored 66.4%, and Sonnet 5 managed just 10.3%. In plain terms, the cheaper model finished more jobs.
Artificial Analysis' independent run agrees: 63.6% for Sonnet 5.5, 59.6% for Opus 5.5, and 59.1% for OpenAI's GPT-6 Astra. The firm highlighted the model's gains in a post on X:
Claude Sonnet 5.5 (max) makes large strides on Terminal-Bench, sitting among the top models for both Terminal-Bench 4.0 and Terminal-Bench-Science. In Terminal-Bench 4.0 it scores 64%, a 50 point increase over Claude Sonnet 5 (max), and slightly above 60% for Opus 5.5 and GPT-6… pic.twitter.com/7CPrhgmfxe
— Artificial Analysis (@ArtificialAnlys) September 28, 2026
Scores also depend on the effort setting, a dial that makes a model think longer for a better answer—and a bigger bill. Anthropic says Sonnet 5.5 at High effort matches GPT6 Sol on FrontierCode for about a fifth of the cost per task.
On GDPval-AA, which grades real-world professional work across 44 occupations using Elo—the chess-style rating system that ranks relative skill—Sonnet 5.5 scored 1844 to Opus 5.5's 1846, effectively a tie. GPT-6 Sol scored 1487.
Rivals match the price. OpenAI cut GPT-6 Sol to $2 and $10 last week, while GPT-5.6 Terra, its mid-tier, lists at $2 and $12. Sonnet 5.5 and GPT-6 Sol now carry identical list prices—though, as the effort settings above show, matching sticker rates do not guarantee matching per-task costs. Anthropic published no benchmark figures for Terra.
The Catch
The trade-off is verbosity. Sonnet 5.5 is a heavy talker: at max effort it wrote about 193,000 tokens per test task, the most Artificial Analysis has measured and roughly 60% more than Opus 5.5. That came to $7.60 per task, about 50% above Sonnet 5—a result that cuts against Anthropic's claim of up to 30% savings.
Anthropic's savings claims rest on lower settings. At Medium effort, the default in its apps, the company says Sonnet 5.5 beats Sonnet 5's best coding score for less than a tenth of the cost. Artificial Analysis says High effort is the best value. For everyday users, that means near-flagship coding at a fraction of the price, as long as the dial stays low.
Caveats remain. Anthropic's benchmark table is self-reported, and Artificial Analysis tested a pre-release build with a bug that Anthropic expects changed little or slightly understated its scores, which is one reason independent results on the shipping build are worth watching. The company also says Opus 5.5 remains clearly stronger at complex work requiring sustained judgment.
Claude Haiku 5.5, built for high-volume, cost-sensitive applications, is due in the coming weeks and would round out Anthropic's 5.5-generation lineup alongside Opus 5.5 and Sonnet 5.5.