Tech • AI • Robotics • Game

VIDEO
ENFR

Daily Podcast full article

Claude Haiku 5.5 slashes pricing for the API economy

Anthropic’s new small model, Claude Haiku 5.5, is not trying to win the AI news cycle by being the biggest brain in the room. It is trying to be the model developers can afford to call all day: a 1-million-token context window, pricing from $0.10 per million input tokens, adaptive effort controls, and a 90% list-price cut on shorter requests.

Generated October 8, 2026 at 12:14 PM1405 words
AI-generated illustration

The headline: cheaper Claude, bigger context, faster jobs

Anthropic released Claude Haiku 5.5 on October 7, positioning it as the cheapest, fastest and most capable small model it has shipped so far . The basic pitch is simple: keep Claude’s enterprise-friendly ergonomics, add a 1-million-token context window, and price the model low enough for high-volume work such as classification, extraction, routing, summarization, compaction, database querying and subagent calls .

The most attention-grabbing number is the input price. For prompts up to 100,000 tokens, Claude Haiku 5.5 starts at $0.10 per million input tokens and $0.50 per million output tokens . That is one-tenth of Haiku 4.5’s listed $1 input and $5 output rates for the same shorter-prompt band, which is why Anthropic describes the new model as priced 90% lower per token for requests up to 100,000 tokens . VentureBeat’s launch coverage framed the same move as a 90% API price reduction, directly tying it to the pressure from OpenAI’s GPT-6 Luna and cheaper model alternatives .

This is not a universal 90% cut across every request shape. Anthropic splits Haiku 5.5 pricing at the 100,000-token line: above that threshold, input rises to $0.50 per million tokens and output to $2.50 per million tokens . Those longer-prompt rates are still lower than Haiku 4.5, but the savings are closer to 50% than 90% . The distinction matters because the model’s 1-million-token context window invites developers to feed it large histories, repositories and document sets, while the sharpest discount is reserved for shorter jobs .

Why the 100K line matters

The launch is as much a pricing architecture story as a model release. Anthropic says prompts up to 100,000 tokens made up about 90% of requests to the previous Haiku model, which explains why the company focused the deepest cut there . In other words, Haiku 5.5 is optimized economically for the common case: short, frequent, repeatable calls that run in the background of software products.

That is the right zone for many enterprise AI workloads. A customer-support agent may need a fast intent classification before escalating to a stronger model. A legal or finance workflow may need hundreds of targeted extractions from known documents. A coding agent may use a smaller sidekick to summarize files, retrieve facts or prepare context before a larger Sonnet or Opus model attempts the hard step. Anthropic’s own documentation describes Haiku 5.5 as built for “high-volume, latency-sensitive” work including classification, routing, extraction and subagent tasks .

But developers will need to watch context carefully. If an application automatically carries long chat histories, system prompts, tool traces or file dumps into every call, it can cross the 100,000-token boundary and lose the headline rate. The result is not that Haiku 5.5 becomes expensive in absolute terms, but that the promised “$0.10 per million input tokens” no longer describes the full bill. At scale, prompt hygiene becomes product strategy.

The model is also a latency play

Haiku 5.5 is not only cheaper. Anthropic says it is its fastest model at standard speed, though the company notes that Opus models in Fast Mode can run faster . The platform documentation lists Haiku 5.5 as the “fastest” model in Anthropic’s current comparison table, ahead of Sonnet 5.5, Opus 5.5 and Fable 5.1 on comparative latency .

This is crucial because many AI products fail less from raw intelligence than from waiting time. A model that is slightly weaker but responds quickly can be more useful for autocomplete, live support, form understanding, background enrichment and agent orchestration. For developers, the practical question is not “Can the model solve the hardest benchmark?” but “Can I afford to call it ten times inside one user interaction without making the product feel slow?”

Anthropic’s launch page leans into that use case. Asana reported more than a 30% reduction in latency for task completions and up to 2.5 times faster inference per agent turn compared with the model it was using, while Box said Haiku 5.5 scored 11 points higher than Haiku 4.5 at roughly half the latency in early testing . These are customer-reported results rather than neutral lab measurements, but they show where Anthropic wants the model to fit: not as the lone genius, but as the fast worker that handles many small steps.

Effort controls turn cost into a dial

A key product change is adaptive thinking with an effort parameter. Haiku 5.5 is the first Haiku-class model to include adjustable effort settings, letting developers trade off cost, speed and intelligence depending on the task . The docs list adaptive thinking as on by default, with medium as the default effort level .

That turns model selection into a more granular decision. Instead of choosing only between Haiku, Sonnet and Opus, a developer can choose Haiku 5.5 at lower effort for routine extraction, or raise effort for harder computer-use and reasoning tasks. Anthropic’s launch benchmarks show Haiku 5.5 outperforming Haiku 4.5 by wide margins on several vendor-reported tests, including OSWorld 2.1, Terminal-Bench 4.0 and knowledge-work evaluations, while Sonnet 5.5 remains ahead on the most demanding work .

The caveat is important: effort settings complicate price-performance comparisons. A maximum-effort run may deliver better benchmark results but consume more tokens and time. A medium-effort run may be the correct production default even if it produces a less spectacular score. In the emerging economics of AI, “best model” increasingly means “best configuration for this workflow.”

Competitive pressure: GPT-6 Luna, open source and unit economics

VentureBeat reported that Haiku 5.5’s lower-tier input and output prices match OpenAI’s GPT-6 Luna at $0.10 and $0.50 per million tokens, while noting that each provider applies different thresholds, caching rules and tokenization behavior . That comparison captures the broader strategic shift: frontier AI vendors are no longer competing only on maximum intelligence. They are competing on usable latency, context length, integration options and predictable unit cost.

That shift has been forced partly by the rise of cheaper and open-source alternatives. If a company can run an open model for routing, extraction or summarization, a proprietary API must justify itself with reliability, safety tooling, deployment reach and total operating cost. Haiku 5.5 is Anthropic’s answer: not a flagship reasoning monster, but a small model priced aggressively enough to stay in the conversation for bulk workloads.

Availability also matters. Anthropic says Haiku 5.5 is available through the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure, with the model ID claude-haiku-5-5 . That multi-cloud footprint gives enterprise buyers a simpler path to test the model inside existing procurement and compliance channels.

The broader package: cache cuts and API credits

The Haiku 5.5 launch came with two additional economics moves. Anthropic said it is halving Claude Sonnet 5.5 cache-read pricing from $0.20 to $0.10 per million tokens, which it says reduces Sonnet 5.5 costs on most agentic tasks by around 20% . It also announced monthly API credits for Claude Max and Team subscribers: $100 per month for Max 5x users, $200 for Max 20x users, and up to $500 pooled across Team subscribers .

That bundle shows Anthropic trying to convert subscription users into API builders. If Haiku 5.5 is the cheap workhorse and Sonnet 5.5 becomes cheaper for cached agent workflows, Anthropic can push developers toward systems where multiple Claude models cooperate: Haiku for quick loops, Sonnet for heavier reasoning, Opus for the hardest work.

Bottom line

Claude Haiku 5.5 matters because it makes price a first-class product feature. The 90% cut on short requests, the $0.10 input rate, the 1-million-token window and adaptive effort controls are all signals that Anthropic sees the next phase of AI adoption as a throughput battle.

The risk is that the slogan is cleaner than the bill. The best rate applies below 100,000 tokens, tokenization has changed, and long-context agent workflows can behave very differently from simple API examples . Still, the direction is unmistakable: frontier labs are learning that enterprise AI is not won only by being brilliant. It is won by being fast enough, safe enough and cheap enough to run everywhere.

Comments

Be the first to comment.

Sources from the last 72 hours

  1. [1]Claude Haiku 5.5Oct 7, 2026, 2:00 AM
  2. [2]Claude Haiku 5.5 - Claude Platform DocsOct 7, 2026, 2:00 AM
  3. [3]Anthropic launches Claude Haiku 5.5 with 90% API price reduction, matching GPT-6 LunaOct 7, 2026, 8:08 PM

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.