Big model to coordinate, small model for repetitive tasks: how to cut the cost of an AI agent
On October 7, Anthropic launched Claude Haiku 5.5, its smallest model, at $0.10 per million input tokens and $0.50 per million output tokens. For a company running AI agents, the important part is architectural. An agent no longer has to run every step on the same model. The strong model can coordinate, while repetitive, well-defined work can move to a model with a per-token price up to 20 times lower, provided it passes the quality test.
Haiku 5.5 brings the price of high-volume tasks down to $0.10
For requests up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Its predecessor, Haiku 4.5, cost $1 and $5. Above the 100,000-token threshold, the price rises to $0.50 and $2.50. According to Anthropic, around 90% of requests to Haiku 4.5 fell below this threshold, a vendor figure that reflects its own customers.
The new tokenizer, the component that splits text into tokens, uses somewhat more tokens for the same task, so Anthropic estimates an average reduction of around 75% versus Haiku 4.5. On the same day, cache reads (the part of a prompt already processed and reused across requests) for Sonnet 5.5 dropped from $0.20 to $0.10 per million tokens, which, by the company’s estimate, cuts the cost of most agentic tasks on that model by about 20%.
Each step of the workflow can have its own model
Anthropic positions Haiku 5.5 for quick, repetitive work: summaries, context compaction, database queries, classification, and acting as a subagent alongside Sonnet 5.5 and Opus 5.5. For complex agentic coding tasks, Anthropic still recommends Sonnet 5.5 and Opus 5.5. The pattern has appeared in Anthropic’s guide to building agents since December 2024. What is new is the price gap, now large enough to matter in an agent’s budget.
An agent that handles a company’s incoming requests has two kinds of steps:
- For the large model: planning the task, ambiguous cases, exceptions, decisions that affect the customer, and the final synthesis.
- For the small model: classifying the request, extracting fields (name, budget, deadline), summarizing an email thread, routing to the right team, repetitive checks, and fixed-format CRM updates.
The question becomes “which model executes each step”. The reasoning level remains a separate cost decision that comes after this one.
How much the gap matters at volume: a simple calculation
Take a workflow that consumes 1 million input tokens and 200,000 output tokens a day, in requests under 100,000 tokens each. At list prices, with no caching and no batch processing:
- Haiku 5.5: $0.10 for input + $0.10 for output = $0.20 a day.
- Sonnet 5.5: $2 for input + $2 for output = $4 a day.
For this mix, the gap is 20 times and grows in proportion to volume. The calculation has limits: the two models’ tokenizers are “similar”, not identical, and retries, escalations to the large model and a higher effort level all add tokens. Above all, the output is not equivalent: the savings exist only for tasks the small model completes at the accepted quality level. That is why we have argued that cost per task matters more than price per token.
The 100,000-token threshold: when the savings shrink
Above the threshold, Haiku 5.5 becomes 5 times more expensive. In the example above, the daily cost would rise from $0.20 to $1. It stays 4 times below Sonnet 5.5 at list prices, because Sonnet 5.5 and Opus 5.5 carry no long-context surcharge. The advantage shrinks without disappearing.
The detail that matters is how the threshold is counted. According to Anthropic’s pricing documentation, it applies to the entire input of a request, including tokens read from and written to the cache, and each request is priced on its own. The threshold should not be tracked only through the “visible” prompt. For a persistent agent, context and cache can push a request into a different cost class.
This leads to architectural decisions: periodic context compaction, memory kept outside the context and retrieved only when relevant, long documents split into chunks, a clean context for each subagent task. Track the distribution of request sizes as well as the average.
How to decide what moves to the small model
Move a task to the small model if:
- its input and output are well defined;
- ambiguity is low and the rules can be written down;
- quality can be measured automatically, on a set of real cases;
- an error is easy to correct and does not reach the customer directly;
- there is an escalation path to the large model when validation fails.
A low token price, on its own, does not justify the move. The useful metric becomes cost per accepted workflow: the total cost of every model involved, including retries, escalations and human review, divided by the number of workflows whose result was accepted. This is how model routing with a clear owner becomes measurable. For an operational AI agent, the map of which model handles which step is set at design time and revisited with every new model release.
Sources: ↗ Anthropic — Introducing Claude Haiku 5.5 · ↗ Claude Platform — Pricing · ↗ Anthropic — Building effective agents · ↗ SiliconANGLE — Haiku 5.5 and Sonnet 5.5 cache pricing
See how we build MassAI agents →