⚡ Uncle Cat AI Radar
ModelsAgentsIndustry

Anthropic Launches Claude Haiku 5.5 at Sharp Price Cut

Claude Haiku 5.5 targets high-volume workloads with faster responses, adjustable effort and operating costs roughly 75% below its predecessor.

What happened

Anthropic released Claude Haiku 5.5, describing it as its fastest and most capable small model. The model is available immediately through Anthropic and the major cloud platforms, including Amazon Web Services, Google Cloud and Microsoft Azure. It is aimed at repetitive, latency-sensitive workloads such as summarization, classification, database queries, customer support and coding subagents.

Anthropic says Haiku 5.5 costs about 75% less to run than Haiku 4.5 on average. The published price is $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, with higher rates above that threshold. The company also introduced an adjustable effort setting, allowing developers to trade speed and reasoning depth against cost.

The same announcement cuts Claude Sonnet 5.5 cache-read pricing in half to $0.10 per million tokens. Anthropic says that reduces Sonnet’s cost by about 20% for many agentic workloads. Early internal testing showed Haiku 5.5 scoring 11 points higher than Haiku 4.5 while operating at roughly half the latency, though the company’s comparisons remain vendor-reported.

Why it matters

The important development is not simply a cheaper small model. Anthropic is pricing the model for the enormous volume of short calls generated by agents, subagents and software automation. Lower inference costs make it easier to keep agents running continuously, but they also intensify pressure on rival providers and on the economics of model hosting. Haiku 5.5 may therefore matter less as a standalone benchmark winner than as another step toward making model calls cheap enough to become routine infrastructure.

Sources