What is a tokenizer — why the same sentence costs 30% more tokens on another model
A tokenizer is the component that cuts text into 'tokens' before an AI language model reads it. The model sees token IDs, not letters, and API prices and context limits are counted in tokens. Each company builds its own token vocabulary from its training data, so the same sentence produces different token counts on different models. In October 2026, Anthropic's Claude Haiku 5.5 cut its per-token price by 90%, but its new tokenizer splits text into about 30% more tokens than Haiku 4.5, so Anthropic's own average saving is about 75%. Languages underrepresented in training data, such as Korean, have fewer whole words in the vocabulary and tend to need more tokens than English to say the same thing
The three lines
- Definition — splits text into tokens the model reads; pricing and context limits are counted in tokens
- Differences — every model family has its own vocabulary; Haiku 5.5 uses ~30% more tokens than 4.5
- Languages — non-English text often needs more tokens; GPT-4o cut a Korean sample 1.7x
Key questions
- What is a tokenizer in AI
- **It converts text into numbered pieces a model can read.** | Step | What happens | Example | |---|---|---| | 1. Split | Break text into vocabulary pieces | unbelievable → un / believ / able | | 2. Number | Map each piece to an ID | 403 / 21893 / 481 (illustrative) | | 3. Read | Model processes the ID sequence | — | | 4. Decode | Turn output IDs back into text | — |
- How many tokens is a word
- **In English, roughly 1 token ≈ 4 characters ≈ 0.75 words (OpenAI's rule of thumb). Other languages vary widely.** | Item | English | Korean and many other languages | |---|---|---| | Rule of thumb | 1 token ≈ 4 characters | No fixed ratio; measure per model | | Pattern | Common words are one token | Often split into syllables or bytes | | Improvement | — | GPT-4o: a Korean sample went from 72 to 41 tokens |
- Does changing tokenizer change API cost
- **Yes — more tokens eat into a per-token price cut.** | Example | Old model | New model | |---|---|---| | Tokens for the same text | 100 | 130 (+30%) | | Price per million tokens | $1.00 | $0.10 | | Relative cost | 100 | 13 |
AI models don't read letters. They read the numbered pieces a tokenizer hands them, which is why the same question can cost different amounts on different models. When Anthropic released Claude Haiku 5.5 on October 7, 2026, it cut per-token prices by 90% but advertised an average saving of 75%, because its new tokenizer cuts the same text into about 30% more tokens (see "Claude Haiku 5.5 pricing"). This explainer covers the device that makes tokens.
1. What a tokenizer does
Language models work only with numbers. Before text goes in, it is split into entries from a fixed "piece vocabulary," and each piece gets an ID. The tokenizer does this, and reverses it when the model answers.
| Stage | What happens | Note |
|---|---|---|
| Build vocabulary | Collect frequent character combinations from a large corpus | Once, before model training |
| Encode | Split input into vocabulary pieces | "unbelievable" → "un" + "believ" + "able" (illustrative) |
| Map to IDs | Assign each piece its vocabulary number | The model reads this sequence |
| Decode | Turn output IDs back into text | Produces the answer |
Why not single characters or whole words? Characters make sequences long and expensive to process; whole words make the vocabulary endless and can't handle unseen words. "Subword" units in between became the standard.
2. How vocabularies are built: BPE
The most common method is byte-pair encoding (BPE), brought to machine translation by Sennrich and colleagues in 2016 and used in variants by most language models, including the GPT family.
| Step | What BPE does |
|---|---|
| 1. Start | Split all text into the smallest units (characters or bytes) |
| 2. Count | Find the most frequent adjacent pair |
| 3. Merge | Add that pair to the vocabulary as one piece |
| 4. Repeat | Continue until the vocabulary hits a target size (e.g., 100K or 200K) |
Frequent strings end up as single pieces; rare ones stay split. Google's SentencePiece, which works on raw text without relying on spaces, is another common tool.
| Tokenizer (examples) | Vocabulary | Used in |
|---|---|---|
| GPT-2 byte-level BPE | ~50K | GPT-2 |
| cl100k_base | ~100K | GPT-3.5, GPT-4 |
| o200k_base | ~200K | GPT-4o onward |
| SentencePiece | 32K (Llama 2) | Llama 2 and others |
Different vocabularies split the same sentence differently. A larger vocabulary packs more text into each token but gives the model more entries to learn; every company strikes its own balance.
3. What changes when the tokenizer changes
| Effect | Detail | Real example |
|---|---|---|
| Cost | APIs bill per token; more tokens shrink a price cut | Haiku 5.5: price −90%, average cost −75% |
| Context limits | "1 million tokens" holds different amounts of text per model | 30% more tokens fills the window sooner |
| Language cost | Underrepresented languages split finer and cost more | GPT-4o: Korean sample 72 → 41 tokens |
| Comparison trap | Per-million-token prices aren't directly comparable across models | Same job, different token counts |
The arithmetic: if text is 100 tokens on the old model and 130 on the new one, and the price falls from $1 to $0.10 per million, cost falls from 100 to 13 — an 87% cut, not 90%. For Haiku 5.5, requests over 100,000 tokens (billed at five times the base rate) pull the blended figure to about 75%.
Korean has long paid this tax. Vocabularies built mostly from English contain few whole Korean words, so words split into syllables or bytes. When OpenAI launched GPT-4o in 2024 with a ~200,000-entry vocabulary, it said a sample Korean sentence dropped from 72 to 41 tokens, 1.7 times fewer. The penalty shrank but didn't vanish. Some researchers want to drop tokenizers entirely and have models read raw bytes, which handles any language and typos but makes sequences longer.
4. FAQ
| Question | Answer |
|---|---|
| How do I count my tokens? | OpenAI offers the tiktoken library and a web tokenizer; Anthropic offers a token-counting API. Measure per model |
| Do spaces and line breaks count? | Yes, they become part of tokens |
| Emoji? | Often several tokens each, because they span multiple bytes |
| Can a tokenizer be swapped on an existing model? | Generally no; the model must be retrained, so tokenizers change with new model generations |
| Are more tokens better? | No. Token count reflects how text is split, not answer quality |
5. What remains unclear
- Haiku 5.5 on non-English text: unpublished; the 30% figure is English-centric. Measure on your own data.
- Measurement differences: 30% (DataCamp) versus 25% (36Kr).
- Byte-level research: the reported Nature paper was confirmed only via an aggregator.
Sources
- Sennrich et al. — Neural Machine Translation of Rare Words with Subword Units (ACL 2016)
- OpenAI — tiktoken (GitHub)
- OpenAI — Hello GPT-4o (language tokenization)
- OpenAI Help — What are tokens and how to count them?
- Hugging Face — Summary of the tokenizers
- DataCamp — Claude Haiku 5.5: Features, Benchmarks, and Pricing