Skip to content
TEN Brief Ten verified stories a day 2026.10.09 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 3 min read · Explainer

What is a tokenizer — why the same sentence costs 30% more tokens on another model

A tokenizer is the component that cuts text into 'tokens' before an AI language model reads it. The model sees token IDs, not letters, and API prices and context limits are counted in tokens. Each company builds its own token vocabulary from its training data, so the same sentence produces different token counts on different models. In October 2026, Anthropic's Claude Haiku 5.5 cut its per-token price by 90%, but its new tokenizer splits text into about 30% more tokens than Haiku 4.5, so Anthropic's own average saving is about 75%. Languages underrepresented in training data, such as Korean, have fewer whole words in the vocabulary and tend to need more tokens than English to say the same thing

Hands sorting colorful wooden pieces of different sizes into rows on a sunny desk

The three lines

  • Definition — splits text into tokens the model reads; pricing and context limits are counted in tokens
  • Differences — every model family has its own vocabulary; Haiku 5.5 uses ~30% more tokens than 4.5
  • Languages — non-English text often needs more tokens; GPT-4o cut a Korean sample 1.7x

Key questions

What is a tokenizer in AI
**It converts text into numbered pieces a model can read.** | Step | What happens | Example | |---|---|---| | 1. Split | Break text into vocabulary pieces | unbelievable → un / believ / able | | 2. Number | Map each piece to an ID | 403 / 21893 / 481 (illustrative) | | 3. Read | Model processes the ID sequence | — | | 4. Decode | Turn output IDs back into text | — |
How many tokens is a word
**In English, roughly 1 token ≈ 4 characters ≈ 0.75 words (OpenAI's rule of thumb). Other languages vary widely.** | Item | English | Korean and many other languages | |---|---|---| | Rule of thumb | 1 token ≈ 4 characters | No fixed ratio; measure per model | | Pattern | Common words are one token | Often split into syllables or bytes | | Improvement | — | GPT-4o: a Korean sample went from 72 to 41 tokens |
Does changing tokenizer change API cost
**Yes — more tokens eat into a per-token price cut.** | Example | Old model | New model | |---|---|---| | Tokens for the same text | 100 | 130 (+30%) | | Price per million tokens | $1.00 | $0.10 | | Relative cost | 100 | 13 |

AI models don't read letters. They read the numbered pieces a tokenizer hands them, which is why the same question can cost different amounts on different models. When Anthropic released Claude Haiku 5.5 on October 7, 2026, it cut per-token prices by 90% but advertised an average saving of 75%, because its new tokenizer cuts the same text into about 30% more tokens (see "Claude Haiku 5.5 pricing"). This explainer covers the device that makes tokens.

1. What a tokenizer does

Language models work only with numbers. Before text goes in, it is split into entries from a fixed "piece vocabulary," and each piece gets an ID. The tokenizer does this, and reverses it when the model answers.

StageWhat happensNote
Build vocabularyCollect frequent character combinations from a large corpusOnce, before model training
EncodeSplit input into vocabulary pieces"unbelievable" → "un" + "believ" + "able" (illustrative)
Map to IDsAssign each piece its vocabulary numberThe model reads this sequence
DecodeTurn output IDs back into textProduces the answer

Why not single characters or whole words? Characters make sequences long and expensive to process; whole words make the vocabulary endless and can't handle unseen words. "Subword" units in between became the standard.

2. How vocabularies are built: BPE

The most common method is byte-pair encoding (BPE), brought to machine translation by Sennrich and colleagues in 2016 and used in variants by most language models, including the GPT family.

StepWhat BPE does
1. StartSplit all text into the smallest units (characters or bytes)
2. CountFind the most frequent adjacent pair
3. MergeAdd that pair to the vocabulary as one piece
4. RepeatContinue until the vocabulary hits a target size (e.g., 100K or 200K)

Frequent strings end up as single pieces; rare ones stay split. Google's SentencePiece, which works on raw text without relying on spaces, is another common tool.

Tokenizer (examples)VocabularyUsed in
GPT-2 byte-level BPE~50KGPT-2
cl100k_base~100KGPT-3.5, GPT-4
o200k_base~200KGPT-4o onward
SentencePiece32K (Llama 2)Llama 2 and others

Different vocabularies split the same sentence differently. A larger vocabulary packs more text into each token but gives the model more entries to learn; every company strikes its own balance.

3. What changes when the tokenizer changes

EffectDetailReal example
CostAPIs bill per token; more tokens shrink a price cutHaiku 5.5: price −90%, average cost −75%
Context limits"1 million tokens" holds different amounts of text per model30% more tokens fills the window sooner
Language costUnderrepresented languages split finer and cost moreGPT-4o: Korean sample 72 → 41 tokens
Comparison trapPer-million-token prices aren't directly comparable across modelsSame job, different token counts

The arithmetic: if text is 100 tokens on the old model and 130 on the new one, and the price falls from $1 to $0.10 per million, cost falls from 100 to 13 — an 87% cut, not 90%. For Haiku 5.5, requests over 100,000 tokens (billed at five times the base rate) pull the blended figure to about 75%.

Korean has long paid this tax. Vocabularies built mostly from English contain few whole Korean words, so words split into syllables or bytes. When OpenAI launched GPT-4o in 2024 with a ~200,000-entry vocabulary, it said a sample Korean sentence dropped from 72 to 41 tokens, 1.7 times fewer. The penalty shrank but didn't vanish. Some researchers want to drop tokenizers entirely and have models read raw bytes, which handles any language and typos but makes sequences longer.

4. FAQ

QuestionAnswer
How do I count my tokens?OpenAI offers the tiktoken library and a web tokenizer; Anthropic offers a token-counting API. Measure per model
Do spaces and line breaks count?Yes, they become part of tokens
Emoji?Often several tokens each, because they span multiple bytes
Can a tokenizer be swapped on an existing model?Generally no; the model must be retrained, so tokenizers change with new model generations
Are more tokens better?No. Token count reflects how text is split, not answer quality

5. What remains unclear

  • Haiku 5.5 on non-English text: unpublished; the 30% figure is English-centric. Measure on your own data.
  • Measurement differences: 30% (DataCamp) versus 25% (36Kr).
  • Byte-level research: the reported Nature paper was confirmed only via an aggregator.

Sources

  1. Sennrich et al. — Neural Machine Translation of Rare Words with Subword Units (ACL 2016)
  2. OpenAI — tiktoken (GitHub)
  3. OpenAI — Hello GPT-4o (language tokenization)
  4. OpenAI Help — What are tokens and how to count them?
  5. Hugging Face — Summary of the tokenizers
  6. DataCamp — Claude Haiku 5.5: Features, Benchmarks, and Pricing

Verification

Published
Last modified
Cross-check
Checked against 6 independent sources.
Unverified
  • Haiku 5.5's extra token count is put at about 30% (DataCamp) and about 25% (36Kr); figures for non-English text are unpublished.
  • GPT-4o's 1.7x reduction for Korean (72 to 41 tokens) was measured by OpenAI on a single sample sentence.
  • A reported Nature paper on 'byteification' (converting subword models to byte-level for under 1% of a pretraining budget, about 49 billion tokens) was confirmed only through an aggregator.
  • Token splits and IDs in examples are illustrative, not actual output from a specific model.
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related