⚡ Uncle Cat AI Radar
ModelsAgentsIndustry

Anthropic Releases Faster, Cheaper Claude Sonnet 5.5

Anthropic says Sonnet 5.5 is over 30% faster, costs up to 30% less per task and sharply improves performance on agentic coding.

A stronger middle tier

Anthropic released Claude Sonnet 5.5, positioning it between the company’s more capable Opus 5.5 and its forthcoming high-volume Haiku 5.5. The new model keeps Sonnet 5’s list price of $2 per million input tokens and $10 per million output tokens, while Anthropic says it is more than 30% faster and can cost up to 30% less per task because it uses fewer tokens.

The largest reported gain is on Terminal-Bench 4.0, an agentic coding evaluation. Sonnet 5.5 scored 70.6%, compared with 10.3% for Sonnet 5 in Anthropic’s published results. The company also says the model is close to Opus 5.5 on real-world work evaluations and can handle long-horizon tasks and image-based reasoning.

Distribution is part of the launch

Sonnet 5.5 became available through Anthropic’s platform and was quickly integrated into products including Claude Code, Devin, Cursor, GitHub Copilot, AWS and other developer environments. That distribution gives the model immediate access to the workflows where speed, reliability and per-task cost matter more than a model’s headline benchmark position.

Anthropic says the model launches with cyber safeguards and fallback mechanisms comparable to those used for more capable systems. The company describes the safeguards as targeted, leaving ordinary software development and most life-science work unaffected.

Why it matters

Sonnet 5.5 shows how model competition is shifting toward useful work completed per unit of time and money. A model that approaches a flagship system while generating faster and consuming fewer tokens could alter which models developers choose as defaults. The unresolved question is whether the benchmark gains persist across messy production tasks, where tool reliability, context management and recovery from errors often determine the result.

Uncle Cat take: The striking figure is Terminal-Bench’s 70.6% versus 10.3%, but production buyers will judge whether that leap survives outside Anthropic’s test harness.

Sources