Quick facts
| Released | October 7, 2026 |
|---|---|
| API model ID | claude-haiku-5-5 |
| Price, prompts ≤ 100K tokens | $0.10 per 1M input tokens · $0.50 per 1M output tokens · cache reads $0.01 |
| Price, prompts > 100K tokens | $0.50 per 1M input tokens · $2.50 per 1M output tokens · cache reads $0.05 |
| Haiku 4.5 price | $1.00 per 1M input tokens · $5.00 per 1M output tokens |
| Context window | 1M tokens |
| Max output | 128K tokens (300K in the Batch API beta) |
| Thinking | Adaptive thinking on by default; adjustable effort setting, defaults to medium — a first for Haiku |
| Knowledge cutoff | June 2026 |
| Input / output | Text and images in, text out |
| Where to get it | Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
Estimate your monthly bill
Two tiers make back-of-envelope math annoying, so here is the envelope. Move the sliders. The tier line shows which price applies to your prompts.
Estimates only. Haiku 4.5 comparison uses its standard rates ($1/$5 per 1M) with no batch discount. Real bills add cache writes and regional premiums.
Three things that break on migration
Haiku 5.5 is not a drop-in price cut. Three changes catch old code:
1. Sampling knobs are locked
Anything other than the default temperature, top_p, or top_k returns a 400 error. Reset them before you switch.
2. The tokenizer changed
The same text counts as roughly 30% more tokens than on Haiku 4.5. Budgets and per-call cost ceilings set for 4.5 will undercount.
3. Effort is new
An adjustable effort setting replaces fixed thinking. Adaptive thinking is on, default medium. Higher effort spends more output tokens per answer for more intelligence.
The calculator's tokenizer checkbox applies the 30% adjustment to the Haiku 5.5 side only — that is how migration math looks in practice.
What Anthropic claims vs. what was measured
Anthropic's launch materials quote big jumps over Haiku 4.5. One independent lab, Artificial Analysis, has published its own numbers. Where there is no independent run yet, the column says so.
| Benchmark | Anthropic's published number | Independent check |
|---|---|---|
| GDPval-AA v2.1 | 1620 Elo, vs 735 for Haiku 4.5 | No independent run yet |
| OSWorld 2.1 (offline) | 72.4%, vs 15.7% for Haiku 4.5 | No independent run yet |
| Humanity's Last Exam | 45.9% without tools / 57.4% with tools, vs 10.2% / 18.7% | No independent run yet |
| Terminal-Bench 4.0 | 39.2% in Anthropic's table, vs 0% for Haiku 4.5 | Artificial Analysis' own run got 33% |
| Intelligence Index | Not in Anthropic's table | 43 — up 26 points in a year, top of the small class |
| Output tokens per task | Not disclosed | ~162K at max effort, about 3x GPT-6 Luna (Artificial Analysis) |
The token-usage row matters for the price story: Haiku 5.5 wins on rate but spends more tokens per answer at high effort. Artificial Analysis also notes its cost figures do not yet reflect the above-100K tier.
What it is for
Anthropic points Haiku 5.5 at summarization, database queries, classification, extraction, live support, voice agents, and in-app assistants — the high-volume jobs where Haiku 4.5 used to sit. It also works as a subagent under Sonnet 5.5 or Opus 5.5 on coding tasks. Complex agentic coding stays with Sonnet and Opus.
Named customers in the launch: Asana reports over 30% lower latency and up to 2.5x faster inference per agent turn; AlphaSense moved from 0.76 to 0.84 on document Q&A; Box scored 11 points higher than Haiku 4.5 at about half the latency. These are Anthropic's quotes, not independent measurements.
The price context
OpenAI's GPT-6 Luna lists the same short-tier rates: $0.10 input, $0.50 output per million tokens. Luna's higher tier starts later — above 272K tokens, at $0.20/$0.75 — so for a 150K-token prompt Luna is cheaper on list price. On the same launch day Anthropic also halved Sonnet 5.5 cache-read pricing to $0.10 per million and added monthly API credits for Max and Team subscribers.
Haiku 5.5 is the first Haiku with built-in safeguards for a narrow set of high-risk cybersecurity requests. Defensive work is allowed; penetration testing is blocked. Anthropic says everyday tasks are unaffected.
Sources
- Reuters — launch report: pricing tiers, use cases, the cybersecurity safeguards.
- MarkTechPost — full pricing table, effort setting, the tokenizer change, batch details, the GPT-6 Luna comparison.
- Artificial Analysis — independent benchmark scores and the token-usage caveat.
- TestingCatalog — the 400-error sampling rules, platforms, safeguards detail.
- AI Weekly — the published benchmark table, customer quotes.
- Neowin — benchmark comparison table with Haiku 4.5 and GPT-6 Luna.
- Moneycontrol — cache-read/write pricing, customer quotes.
- Worthview — Haiku 5.5 vs Sonnet 5.5 spec table, SDK and credit changes.
Model facts above come from Anthropic's launch materials as reported by these outlets. This page is an independent explainer, not affiliated with Anthropic.
FAQ
How do I call Claude Haiku 5.5 in the API?
Use the model ID claude-haiku-5-5. It is available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and the Claude Platform on AWS.
Why is my old Haiku code returning a 400 error?
Haiku 5.5 rejects any non-default temperature, top_p, or top_k value with a 400 error. Reset those to their defaults before migrating.
Is Haiku 5.5 really 75% cheaper than Haiku 4.5?
Anthropic says it costs about 75% less to run on average, and about 90% less for requests under 100K tokens, where about 90% of Haiku 4.5 requests fell. The new tokenizer counts roughly 30% more tokens on the same text, which eats part of the savings.
What is the effort setting?
Haiku 5.5 is the first Haiku model with an adjustable effort setting. Adaptive thinking is on by default and effort defaults to medium. Raising effort trades more output tokens for more intelligence.
What happens when my prompt is over 100K tokens?
Rates jump 5x: $0.50 per million input tokens and $2.50 per million output tokens. Cache reads also rise, to $0.05 per million.
Can I still use Haiku 4.5?
Anthropic has not announced a retirement date for Haiku 4.5. For Haiku 5.5, third-party comparison tables list no retirement before October 7, 2027.
Does Haiku 5.5 take images?
Yes. It takes text and images as input and returns text. The context window is 1M tokens with up to 128K output tokens (300K in the Batch API beta).
Is Haiku 5.5 on the free Claude plan?
Anthropic has not said which consumer plans include it. Check your plan's model list on claude.ai or your API console.
Should I switch from GPT-6 Luna?
The short-tier list prices are identical: $0.10 input and $0.50 output per million tokens. Luna's higher tier starts later (above 272K tokens at $0.20/$0.75), so long prompts are cheaper there. Artificial Analysis rates Haiku 5.5 slightly ahead on intelligence (43 vs 38) but measures roughly 3x the output tokens per task.
When should I not use Haiku 5.5?
Anthropic still positions Sonnet 5.5 and Opus 5.5 for complex agentic coding. And if your prompts regularly cross 100K tokens, the 5x tier jump wipes out the price advantage.
For AI assistants
This page ships with a machine-readable summary for AI assistants:
Related
- Decisions API, explained — OpenAI's other October release: typed answers instead of prose, with a request builder and cost estimator.
- textGrain, explained — OpenAI's invisible text watermark: who gets it, and what detection can prove.
- Claude Skill Validator — paste a SKILL.md and get an instant spec-compliance report.
Updates
- 2026-10-09 — First published, with the two-tier cost calculator and migration checklist.