What happened?

Anthropic released Claude Haiku 5.5 on 7 October. It is the fast, lower-cost member of the Claude family, built for work that happens thousands of times: classify a request, find a fact in a document, summarise a conversation, or help a larger agent with one small part of a job.

The headline is the price. On the Claude API, prompts up to 100,000 tokens cost $0.10 per million input tokens and $0.50 per million output tokens. That is 90% below Haiku 4.5's token rates. Above that prompt size, Haiku 5.5 costs $0.50 input and $2.50 output. Those rates are still lower than its predecessor's, but they are not the headline bargain.

USD per million tokens Prompt up to 100k Prompt over 100k
Input $0.10 $0.50
Output $0.50 $2.50
Cache read $0.01 $0.05

Anthropic estimates an average running-cost reduction of about 75%, not 90%. The distinction matters: price per token is not price per completed task. Its migration guide says the newer tokenizer counts approximately 30% more tokens for the same text, with the exact increase depending on the content.

Haiku also gains adjustable reasoning effort, a 1 million-token context window, and up to 128,000 output tokens. Adaptive thinking is enabled by default, at medium effort. Developers can call it now as claude-haiku-5-5.

Why this is interesting

The promising role is not "put the cheapest model in charge of everything." It is give the small model a clear job.

An incoming support message might need a category, a relevant policy lookup, and a short draft. A document workflow might need dates, amounts, and a summary before a person reviews it. A larger coding agent might need a helper to find the right function or summarise a file. These are possible applications, not results from our own deployments.

Anthropic's early customer reports give that idea some substance. AlphaSense tested 400 document questions and reported a score of 0.84 versus 0.76 for Haiku 4.5. Asana reported over 30% lower task-completion latency and up to 2.5 times faster inference per agent turn against its existing model. Those are different measurements, on the customers' own suites, quoted in Anthropic's announcement - not a universal speed guarantee.

The useful design pattern is a division of labour: Haiku handles bounded, repetitive steps; Sonnet or Opus handles the harder synthesis and planning. A person approves consequential actions. Lower token costs make that pattern more attractive to test, especially when the same short request runs all day.

Faster still needs the right settings

Haiku 5.5 is the first Haiku with an effort control. Independent developer Simon Willison's SVG test makes the trade-off tangible: his low-effort pelican took 7 seconds; the max-effort version took 5 minutes 9 seconds. It is one creative prompt, not a business benchmark, but it illustrates why a fast model can still produce a slow answer when asked to think extensively.

Artificial Analysis separately reports an Intelligence Index score of 43 at max effort, versus 29 at low effort. Those are results for different settings of the same model, not a reason to run every classification request at max.

There are limits to the headline capability numbers, too. Anthropic's OSWorld 2.1 offline result is 72.4% partial credit, while the strict full-task pass rate is 37.1%. That evaluation used 82 offline computer tasks at max effort. It shows a substantial capability improvement, not a 72.4% guarantee that an agent will finish your workflow.

Before you switch

The migration guide is worth reading before changing the model ID. Old thinking budgets must move to adaptive thinking. Sampling parameters such as temperature, top_p, and top_k should be removed. Applications must read response blocks by type, leave room for thinking within the output limit, and handle refusals explicitly.

A sensible first trial is one repetitive workload with examples you can grade. Measure accepted results, response time, retries, and the full bill. Give the helper only the access it needs, and route uncertain or consequential cases to a stronger model or a person. The cheapest first attempt is not always the cheapest successful result.

Bottom line

Haiku 5.5 makes the routine layer of an agent system much more affordable. Alongside it, Anthropic has halved Sonnet 5.5 cache-read pricing to $0.10 per million tokens and announced monthly API credits rolling out to Max and Team subscribers.

For businesses, the opportunity is straightforward: try automating a well-defined piece of repetitive work that previously felt too expensive to run at volume. Start small, check the results, and keep human judgement where it belongs.