Skip to content
TEN Brief Ten verified stories a day 2026.10.09 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 3 min read · Breaking

Claude Haiku 5.5 price is $0.10 per million tokens — why a 90% cut averages 75%

Anthropic released its small model Claude Haiku 5.5 on October 7, 2026. For requests up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, 90% less than Haiku 4.5 ($1 and $5). But once a request passes 100,000 tokens, the entire request is billed at five times that rate ($0.50 and $2.50), and a new tokenizer splits the same text into about 30% more tokens. That is why Anthropic's own average saving is about 75%, not 90%. Haiku 5.5 scored 72.4% on the OSWorld 2.1 computer-use test, 4.6 times Haiku 4.5's 15.7%, has a 1 million-token context window, and is available on every Claude plan including Free

Developer seen from behind at a standing desk with two monitors in a sunny startup office

The three lines

  • Price — $0.10 in / $0.50 out per million tokens up to 100K tokens, 90% below Haiku 4.5
  • Catch — over 100K, the whole request costs 5x; ~30% more tokens; average saving ~75%
  • Performance — OSWorld 2.1 72.4% (4.5: 15.7%), 1M context, supported until at least Oct 7, 2027

Key questions

Claude Haiku 5.5 pricing
**Two tiers split at 100,000 tokens per request (per million tokens).** | Rate | ≤100K | >100K | Haiku 4.5 | |---|---|---|---| | Input | $0.10 | $0.50 | $1.00 | | Output | $0.50 | $2.50 | $5.00 | | Cache read | $0.01 | $0.05 | $0.10 | | Cache write (5 min) | $0.125 | $0.625 | $1.25 | Batch API: a further 50% off.
Haiku 5.5 100K token price cliff
**The whole request moves to the higher tier, not just the tokens above 100K.** | Request size | Input rate applied | |---|---| | 99,000 tokens | $0.10 | | 101,000 tokens | $0.50 for the entire request | Long-document analysis, chatbots that resend full history and codebase agents hit this tier most.
Claude Haiku 5.5 vs GPT-6 Luna
**Same API price; Haiku 5.5 led on Anthropic's published benchmarks.** | Item | Haiku 5.5 | GPT-6 Luna | |---|---|---| | Price per 1M (in/out) | $0.10 / $0.50 | $0.10 / $0.50 | | OSWorld 2.1 | 72.4% | 48.9% | | Terminal-Bench 4.0 | 39.2% | 16.4% | | GDPval-AA v2.1 (Elo) | 1620 | 1437 |

The headline cut is 90%. The number Anthropic itself advertises is 75%. The 15-point gap explains the whole pricing design. Anthropic released Claude Haiku 5.5 on October 7, 2026, two weeks after Opus 5.5 and nine days after Sonnet 5.5. It targets high-volume, cost-sensitive work such as summarizing, classification and database lookups. The same day, OpenAI opened GPT-6 Luna to free ChatGPT users (see "GPT-6 free on ChatGPT"), putting the two companies' small models at exactly the same API price.

1. The price sheet: two tiers split at 100,000 tokens

Per million tokensHaiku 5.5 (request ≤100K)Haiku 5.5 (>100K)Haiku 4.5
Input$0.10$0.50$1.00
Output$0.50$2.50$5.00
Cache read$0.01$0.05$0.10
Cache write (5 min)$0.125$0.625$1.25
vs Haiku 4.590% cheaper50% cheaper—

The Batch API takes another 50% off both tiers, and thinking tokens bill as output. The key detail is that the 100,000-token line applies per request. A 99,000-token request pays $0.10 per million input tokens; a 101,000-token request pays $0.50 on every token, not just the extra 1,000. Long contract analysis, support bots that resend full conversation history, and agents that hold a large codebase in context are most exposed. Anthropic says 90% of requests to its previous model were under 100,000 tokens.

2. Why the average is 75%: about 30% more tokens

Haiku 5.5 shares a new tokenizer with Sonnet 5.5 and Opus 5.5 that splits the same text into roughly 30% more tokens than Haiku 4.5 (DataCamp; 36Kr estimates about 25%). A 90% cut in the per-token price shrinks less once token counts rise.

Worked example (≤100K tier)Haiku 4.5Haiku 5.5
Tokens for the same text (assumed)100130
Input price per million$1.00$0.10
Relative cost10013 (87% saving)

Mix in some requests above 100,000 tokens, where the saving is only 50%, and the blended figure drops further — hence Anthropic's "about 75% cheaper on average." One hands-on test by DataCamp's author had both models judge 24 invoices: both scored 24 out of 24, and the full run cost $0.0098 on Haiku 5.5 against $0.25 on Haiku 4.5, about one twenty-fifth.

3. Performance: a different class from 4.5, below Sonnet 5.5

Figures from Anthropic's launch post:

BenchmarkHaiku 5.5Haiku 4.5Sonnet 5.5GPT-6 Luna
OSWorld 2.1 (computer use)72.4%15.7%83.9%48.9%
Terminal-Bench 4.039.2%0.0%70.6%16.4%
GDPval-AA v2.1 (Elo)162073518401437
Humanity's Last Exam (no tools)45.9%10.2%56.9%—
FrontierCode 1.146.4%—52.1%42.4%
SpecDetail
Context1M tokens
Max output128K tokens (300K via Batch)
Effort settingFirst Haiku model with low-to-max effort
PlatformsClaude API, Amazon Bedrock, Google Cloud, Microsoft
Claude appsAll plans, including Free
SupportNot retired before Oct 7, 2027

Anthropic also halved Sonnet 5.5's cache-read price from $0.20 to $0.10 per million tokens, which it says cuts the cost of most agentic work on Sonnet 5.5 by about 20%.

4. What remains unclear

  • Non-English token counts: how much more (or less) the new tokenizer splits Korean or other languages is unpublished; the 30% figure is English-centric. Teams should measure on their own data.
  • Benchmark verification: the figures are Anthropic's; independent replication wasn't found.
  • Chinese competition: Qwen3.8 Flash ($0.15 in / $0.47 out) and GLM 5.3 Flash ($0.15 / $0.50) sit at similar prices; no head-to-head performance data is available.
  • Related: how tokenizers work is explained in "What a tokenizer is."

Sources

  1. Anthropic — Introducing Claude Haiku 5.5
  2. DataCamp — Claude Haiku 5.5: Features, Benchmarks, and Pricing
  3. OrcaRouter — Claude Haiku 5.5: the 100K Price Cliff Nobody Mentioned
  4. Yahoo Finance — Anthropic reveals Haiku 5.5 model as AI pricing war intensifies
  5. 36Kr — Claude Haiku 5.5 Cuts Prices to Rock-Bottom Levels

Verification

Published
Last modified
Cross-check
Checked against 5 independent sources.
Unverified
  • The tokenizer's extra token count is put at about 30% (DataCamp, citing Anthropic guidance) and about 25% (36Kr). The real figure varies by language and text; no Korean-language figure has been published.
  • Benchmarks are Anthropic's launch figures as reproduced by outlets; independent replication was not found.
  • Anthropic's statement that 90% of requests to the previous model were under 100,000 tokens does not specify the period or customer base.
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related