Model · stat sheet
Claude Haiku 4.5
Cheap and quick for classification, routing and the hundred small calls inside a bigger system.
What it is
Claude Haiku 4.5 is Anthropic's small, fast tier — and note the version number: it is a 4.5-generation model still in the line-up alongside the Claude 5 tiers, not a Haiku 5. It differs from its siblings in ways that matter. The context window is 200k rather than 1M, maximum output is 64k rather than 128k, and it has the older extended-thinking toggle instead of the adaptive thinking available on Sonnet 5 and above.
It is not a cheap frontier model. It is a different job: the hundred small calls inside a bigger system.
How the composite is built
Overall 80 weights Coding 24%, Terminal 20%, Reasoning 20%, Tool use 15%, Context 13%, Speed 8%. The weighting is unkind to this model — it is strongest on the two axes the index cares least about.
- Coding 78
- Below most of the open-weights field. Fine for a well-specified function or a mechanical edit; not the model to hand an under-specified refactor.
- Terminal 81
- Higher than its Coding score, which is unusual and revealing — it follows a defined procedure well even when it would not have designed that procedure itself.
- Reasoning 76
- The weakest axis, and the boundary of what it should be asked to do. Classification and routing, yes; judgement calls with real consequences, no.
- Tool use 85
- The standout. Higher than DeepSeek V4's 83 and Kimi k3's 80, from a model at a fraction of the capability elsewhere. Clean schema-following is what a small model can genuinely be good at.
- Context 72
- The lowest of any Anthropic model here, and the honest consequence of a 200k window against the line's 1M. Plan for it: Haiku is for bounded inputs.
- Speed 94
- The board high, ahead of Grok 5 Mini's 91. This is the product. At 8% weight the composite gives it almost no credit, which is why the overall 80 understates how useful it is in the right slot.
Where it fits
Inside a larger system rather than as the system. Intent classification, routing to a bigger model, extracting structured fields, summarising a tool response before it hits an expensive context, drafting commit messages, triaging logs. Anywhere the call happens thousands of times and each one is small and well-defined.
Limits
- 200k context, not 1M. Code written against the Claude 5 tiers will silently overflow here. This is the single most common migration mistake.
- No adaptive thinking. You get the older extended-thinking toggle, so prompting tricks tuned for the 5 line do not transfer.
- Value is rated Poor at 8.4. 16.9 index points per $2.00/M blended. Cheaper open models beat it comfortably on price-per-point — GPT-5.6 Luna rates 82.9 on the same measure.
- Reasoning 76 fails quietly. A small model rarely announces that a task was beyond it; it returns something plausible. Validate its output rather than trusting it.
Price and access
$1 per million input tokens and $5 per million output, blending to $2.00/M — half of Sonnet 5 and a tenth of Fable 5. Available on the Claude API, OpenRouter and OpenCode Zen. Worth noting that within Anthropic's own line it is the price floor, but it is not cheap against the wider board. Last scored 15 Sep 2026.
Alternatives on this board
- GPT-5.6 Luna — 81 overall at $0.20/$1.20, Speed 90, Context 89. Better score, better context, less than a quarter of the blended price. The obvious comparison.
- Gemini 3.6 Flash — 84 at $0.75/$3.75 with Context 93 and Speed 88. The step up if your small calls need long inputs.
- Mistral Small 3.2 — 79 at $0.09/$0.25 with open weights, if you want the same job done inside your own infrastructure.
- Claude Sonnet 5 — 88 at $2/$10, for when the small call turns out not to be small.
Sources
- Anthropic's model overview — confirm the 200k window, 64k output cap and thinking behaviour before you design around them.
- Haiku 4.5 on OpenRouter for live pricing and providers.
- On this site: how the Claude tiers differ, the price-to-performance guide, and the full model board.
Scores are fullauto.online's composite index (0–100): Coding 24% · Terminal 20% · Reasoning 20% · Tool use 15% · Context 13% · Speed 8%. Editorial, not a vendor benchmark; 2026 tiers are early reads. Last scored 15 Sep 2026 · back to the leaderboard.