
Tech • AI • Robotics • Game
Anthropic has launched Claude Haiku 5.5 as its lowest-cost and fastest small model yet, with sharply lower pricing and major gains on routine office, browser, and automation tasks.
Haiku 5.5 is positioned as the cheapest model in Anthropic’s Claude 5.5 family, built for high-volume and cost-sensitive work. Pricing falls to $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, down from $1 and $5 respectively on the prior Haiku 4.5. That represents a roughly 90% cut on both input and output rates, while Anthropic says average operating cost is about 75% lower overall.
The lowest rate only covers prompts under 100,000 tokens, or roughly 55,000 words. Above that threshold, pricing rises to $0.50 per million input tokens and $2.50 per million output tokens. Even at that higher tier, the model remains cheaper than the previous Haiku generation, and Anthropic says about 90% of requests to the older model were below the threshold.
Compared with Sonnet 5.5, priced at $2 per million input tokens and $10 per million output tokens, Haiku 5.5 is about 20 times cheaper. That gap could make a difference for repetitive jobs such as summarizing documents, classifying emails, tagging support tickets, extracting figures from reports, or handling live customer-service interactions where requests arrive at scale.
The new model shows a major capability jump from older Haiku versions, which had been regarded as weak on many practical tasks. On a computer-use benchmark, Haiku 4.5 scored 15.7%, while Haiku 5.5 reached 72.4%. On knowledge-work tests covering real office tasks, it more than doubled the earlier model and also outperformed a rival small model referenced in the comparison.
On agentic coding, Haiku 4.5 reportedly scored 0%, while Haiku 5.5 reached 39.2%. That is a significant leap, but still well behind Sonnet 5.5, which scored about 70% on the same test. The gap suggests Haiku is suitable for simpler coding and structured automation, while more complex software work still favors larger models.
Haiku 5.5 is the first Haiku model with an adjustable effort setting, allowing users to tune for lower cost or higher reasoning effort. That gives teams a way to balance speed, price, and output quality depending on the task, especially in production systems where the same workflow may need different settings for triage, extraction, or validation.
The model’s strongest use cases are jobs that are narrow, clear, and frequent: inbox sorting, transcript summaries, invoice total extraction, support-ticket tagging, and quick browser-driven actions. Its speed also makes it a fit for live chat and other latency-sensitive applications. It is less suited to tasks with many steps or cases where an error would be costly.
Anthropic says Haiku 5.5 pairs well with Sonnet and Opus as a sub-agent. A typical workflow would let a larger model handle planning or synthesis while Haiku performs lower-cost support tasks such as pulling a figure from a financial report, scraping a page, or classifying inputs. That setup helps reserve expensive reasoning capacity for harder work.
Early access users cited speed gains alongside lower cost. ASA reported more than a 30% reduction in latency compared with its current model, while Box said Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency. Those results reinforce Anthropic’s claim that the model is optimized for fast production use.
Anthropic also lowered Sonnet 5.5 cached reads from $0.20 to $0.10 per million tokens, which it says makes Sonnet about 20% cheaper on many agentic workloads. In addition, Max and Team subscribers now receive monthly API credits: $100 for Max 5x, $200 for Max 20x, and up to $500 shared across eligible team plans.
Haiku 5.5 marks a notable shift in Anthropic’s lineup by making everyday automation far cheaper without leaving capability at older small-model levels. The main trade-off remains clear: for simple, repeated tasks it may be enough, but complex coding and high-stakes reasoning still favor larger models.
Ask a question