Claude Haiku 5.5 price is $0.10 per million tokens — why a 90% cut averages 75%
Anthropic released its small model Claude Haiku 5.5 on October 7, 2026. For requests up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, 90% less than Haiku 4.5 ($1 and $5). But once a request passes 100,000 tokens, the entire request is billed at five times that rate ($0.50 and $2.50), and a new tokenizer splits the same text into about 30% more tokens. That is why Anthropic's own average saving is about 75%, not 90%. Haiku 5.5 scored 72.4% on the OSWorld 2.1 computer-use test, 4.6 times Haiku 4.5's 15.7%, has a 1 million-token context window, and is available on every Claude plan including Free
The three lines
- Price — $0.10 in / $0.50 out per million tokens up to 100K tokens, 90% below Haiku 4.5
- Catch — over 100K, the whole request costs 5x; ~30% more tokens; average saving ~75%
- Performance — OSWorld 2.1 72.4% (4.5: 15.7%), 1M context, supported until at least Oct 7, 2027
Key questions
- Claude Haiku 5.5 pricing
- **Two tiers split at 100,000 tokens per request (per million tokens).** | Rate | ≤100K | >100K | Haiku 4.5 | |---|---|---|---| | Input | $0.10 | $0.50 | $1.00 | | Output | $0.50 | $2.50 | $5.00 | | Cache read | $0.01 | $0.05 | $0.10 | | Cache write (5 min) | $0.125 | $0.625 | $1.25 | Batch API: a further 50% off.
- Haiku 5.5 100K token price cliff
- **The whole request moves to the higher tier, not just the tokens above 100K.** | Request size | Input rate applied | |---|---| | 99,000 tokens | $0.10 | | 101,000 tokens | $0.50 for the entire request | Long-document analysis, chatbots that resend full history and codebase agents hit this tier most.
- Claude Haiku 5.5 vs GPT-6 Luna
- **Same API price; Haiku 5.5 led on Anthropic's published benchmarks.** | Item | Haiku 5.5 | GPT-6 Luna | |---|---|---| | Price per 1M (in/out) | $0.10 / $0.50 | $0.10 / $0.50 | | OSWorld 2.1 | 72.4% | 48.9% | | Terminal-Bench 4.0 | 39.2% | 16.4% | | GDPval-AA v2.1 (Elo) | 1620 | 1437 |
The headline cut is 90%. The number Anthropic itself advertises is 75%. The 15-point gap explains the whole pricing design. Anthropic released Claude Haiku 5.5 on October 7, 2026, two weeks after Opus 5.5 and nine days after Sonnet 5.5. It targets high-volume, cost-sensitive work such as summarizing, classification and database lookups. The same day, OpenAI opened GPT-6 Luna to free ChatGPT users (see "GPT-6 free on ChatGPT"), putting the two companies' small models at exactly the same API price.
1. The price sheet: two tiers split at 100,000 tokens
| Per million tokens | Haiku 5.5 (request ≤100K) | Haiku 5.5 (>100K) | Haiku 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache read | $0.01 | $0.05 | $0.10 |
| Cache write (5 min) | $0.125 | $0.625 | $1.25 |
| vs Haiku 4.5 | 90% cheaper | 50% cheaper | — |
The Batch API takes another 50% off both tiers, and thinking tokens bill as output. The key detail is that the 100,000-token line applies per request. A 99,000-token request pays $0.10 per million input tokens; a 101,000-token request pays $0.50 on every token, not just the extra 1,000. Long contract analysis, support bots that resend full conversation history, and agents that hold a large codebase in context are most exposed. Anthropic says 90% of requests to its previous model were under 100,000 tokens.
2. Why the average is 75%: about 30% more tokens
Haiku 5.5 shares a new tokenizer with Sonnet 5.5 and Opus 5.5 that splits the same text into roughly 30% more tokens than Haiku 4.5 (DataCamp; 36Kr estimates about 25%). A 90% cut in the per-token price shrinks less once token counts rise.
| Worked example (≤100K tier) | Haiku 4.5 | Haiku 5.5 |
|---|---|---|
| Tokens for the same text (assumed) | 100 | 130 |
| Input price per million | $1.00 | $0.10 |
| Relative cost | 100 | 13 (87% saving) |
Mix in some requests above 100,000 tokens, where the saving is only 50%, and the blended figure drops further — hence Anthropic's "about 75% cheaper on average." One hands-on test by DataCamp's author had both models judge 24 invoices: both scored 24 out of 24, and the full run cost $0.0098 on Haiku 5.5 against $0.25 on Haiku 4.5, about one twenty-fifth.
3. Performance: a different class from 4.5, below Sonnet 5.5
Figures from Anthropic's launch post:
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 | GPT-6 Luna |
|---|---|---|---|---|
| OSWorld 2.1 (computer use) | 72.4% | 15.7% | 83.9% | 48.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 70.6% | 16.4% |
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1840 | 1437 |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | 56.9% | — |
| FrontierCode 1.1 | 46.4% | — | 52.1% | 42.4% |
| Spec | Detail |
|---|---|
| Context | 1M tokens |
| Max output | 128K tokens (300K via Batch) |
| Effort setting | First Haiku model with low-to-max effort |
| Platforms | Claude API, Amazon Bedrock, Google Cloud, Microsoft |
| Claude apps | All plans, including Free |
| Support | Not retired before Oct 7, 2027 |
Anthropic also halved Sonnet 5.5's cache-read price from $0.20 to $0.10 per million tokens, which it says cuts the cost of most agentic work on Sonnet 5.5 by about 20%.
4. What remains unclear
- Non-English token counts: how much more (or less) the new tokenizer splits Korean or other languages is unpublished; the 30% figure is English-centric. Teams should measure on their own data.
- Benchmark verification: the figures are Anthropic's; independent replication wasn't found.
- Chinese competition: Qwen3.8 Flash ($0.15 in / $0.47 out) and GLM 5.3 Flash ($0.15 / $0.50) sit at similar prices; no head-to-head performance data is available.
- Related: how tokenizers work is explained in "What a tokenizer is."
Sources
- Anthropic — Introducing Claude Haiku 5.5
- DataCamp — Claude Haiku 5.5: Features, Benchmarks, and Pricing
- OrcaRouter — Claude Haiku 5.5: the 100K Price Cliff Nobody Mentioned
- Yahoo Finance — Anthropic reveals Haiku 5.5 model as AI pricing war intensifies
- 36Kr — Claude Haiku 5.5 Cuts Prices to Rock-Bottom Levels