Theme
Tokenization & Vocabulary
Tokenization is the process of chopping continuous text into discrete units — tokens — which are the smallest units a language model truly "sees" and processes. Models don't understand text directly; they only recognize integer IDs from their vocabulary (vocabulary). What language modeling calls "predicting the next word" is precisely "predicting the next token." Tokenization quality determines what the model can learn, how well it learns, and how much compute and cost it incurs — hence it's often called "the most underrated critical component."
One-line summary: the tokenizer defines the model's input alphabet — it determines "what the model can express" (vocabulary coverage) and "how much each sentence costs" (token count).
1. Why Tokenization Is Needed
A model could theoretically model at the character level, but pure character granularity has three problems:
| Problem | Explanation |
|---|---|
| Sequences get too long | An average word has 4–5 characters, inflating sequence length several-fold and making self-attention O(n²) costs explode |
| Semantic units are fragmented | The 5 characters of "apple" are independent; the model must learn to stitch them together, increasing learning burden |
| Knowledge doesn't cross tasks | Subword-level knowledge (e.g., the grammatical meaning of "ing", "tion") cannot be directly reused by a character-level model |
Tokenization seeks a balance between granularity and coverage: smaller units can cover any text (including rare words) but produce longer sequences with fragmented semantics; larger units preserve semantic integrity but lead to sparser vocabularies and harder new-word coverage. Modern mainstream approaches use subword tokenization: common words as wholes, rare words decomposed into composable subword fragments — the best of both worlds.
2. Tokenization Approaches Compared
| Approach | Basic Unit | Vocabulary Source | Pros | Cons | Typical Use |
|---|---|---|---|---|---|
| Character (char) | Individual characters | Fixed | Tiny vocab, covers everything | Long sequences, no semantics | Fallback for rare languages |
| Word | Full words | Corpus statistics | Complete semantics | Huge vocab, many OOVs, morphological explosion | Traditional NLP |
| BPE | Subword (byte merging) | Corpus-statistic merging | Good coverage, reversible, simple to implement | No linguistic priors, unstable for rare morphologies | GPT, Llama family |
| WordPiece | Subword (probability merging) | Corpus-statistic merging | Better merge criterion (likelihood gain) | Depends on space-pre-tokenized input | BERT, T5 family |
| Unigram | Subword (probability reduction) | Reverse-inferred from large vocab | Controllable vocab size, output can sample by probability | Complex implementation | SentencePiece, LLaMA (later) |
| SentencePiece | Subword + whole-sentence encoding | Unified framework of above methods | No space pre-tokenization, byte-level support, cross-lingual | Many concepts, requires configuration | LLaMA, Mistral, Qwen family |
Two core algorithms — BPE (Byte-Pair Encoding) and WordPiece — differ in their merge criteria: BPE always merges the most frequently co-occurring adjacent pair; WordPiece always merges the adjacent pair that maximizes the language model likelihood gain (like "best single-step tokenization"). Unigram goes in the opposite direction: first build an oversized candidate vocabulary, then iteratively remove words based on EM-computed loss contribution until reaching the target size. All three are based on the insight that "subwords = the solution to balancing vocab coverage and sequence length."
Notably, word-level tokenization hasn't been phased out: for morphologically simple languages (like English with regular word forms), it remains intuitive. But for morphologically rich languages (German compound words, Turkish agglutinative), spaceless languages (Chinese/Japanese/Thai), and high-OOV scenarios, subword tokenization's systematic advantages make it the de facto standard for modern LLMs.
3. BPE Algorithm: Step-by-Step Breakdown
BPE was introduced by Sennrich et al. 2016 from the data compression domain into neural machine translation, and is the mainstream algorithm adopted by the GPT series (byte-level variant: byte-level BPE) and the Llama family. The core idea: repeatedly merge the most frequently co-occurring adjacent token pair in the corpus, until the vocabulary reaches a target size.
text
Input: training corpus, target vocabulary size V
Step 0: Split each word into character sequences and count, e.g.
"low"×5 "lower"×2 "newest"×6 "widest"×3
Step 1: Count all adjacent character pairs' frequencies:
("l","o")=7 ("o","w")=7 ("n","e")=6 ("w","e")=8 ...
Step 2: Merge the highest-frequency pair ("w","e") → new vocab entry "we"
Corpus transforms: lo-w·er → "lo" "we" "r" form (illustrative)
Step 3: Repeat: count new adjacent pair frequencies, continue merging highest
("lo","w")=7 → "low"; ("n","e")=6 → "ne"; ("ne","w")=6 → "new"...
Step 4: Stop when vocab size reaches V, training complete
Encoding new text:
"lowest" → greedy match longest vocab entries:
Start at character level, repeatedly apply learned merge rules:
l-o-w-e-s-t → "low" "est" (if "est" is also in vocab)Key properties:
- Byte-level BPE: starting from GPT-2, input is reduced to UTF-8 bytes before merging, so the vocabulary only depends on 256 byte values. No Unicode text will produce out-of-vocabulary (OOV) tokens — this is the foundation of multilingual robustness.
- Deterministic greedy encoding: encoding applies greedy merges by longest match/learned rules, fast and reproducible.
- Deficiency: rules are independent of language structure — whether morphological fragments like "st" and "tion" form, and in what order, depends entirely on frequency.
A complete example clarifies how byte-level BPE "grows": if the corpus heavily features "low" and "lower", training would first merge "lo" (if it's the highest-frequency pair), then form "low", and finally may form "lower"; while in "newest" and "widest", the high-frequency "est" is merged independently and reused for all words ending in -est. After training, "lowest" would be greedily split into two subwords: "low" + "est" — this is exactly how subword approaches work in practice: common words as wholes, rare words in chunks.
python
# Merge logic for byte-level BPE (pseudocode, illustrating the training loop)
def train_bpe(corpus: dict[str, int], vocab_size: int):
vocab = set() # vocabulary
while len(vocab) < vocab_size:
pairs = count_adjacent_pairs(corpus) # count adjacent token pair frequencies
best = max(pairs, key=pairs.get) # pick the highest-frequency pair
vocab.add(best) # add to vocabulary
corpus = merge_pair(corpus, best) # replace that pair in corpus
return vocabPractical implementation
No need to implement from scratch in production — SentencePiece (Google, open source) unifies BPE/Unigram/WordPiece and eliminates space pre-tokenization; the Hugging Face tokenizers library provides a high-performance Rust implementation. Writing BPE from scratch is about understanding principles, not reinventing the wheel.
4. Special Tokens & Vocabulary Design
The vocabulary isn't just content tokens; it also contains special tokens used by the model's protocol:
| Special Token | Purpose | Common Notation |
|---|---|---|
<bos> | Sequence start | GPT-2 doesn't use; Llama does |
<eos> | Sequence end / dialogue turn boundary | GPT-style uses <endoftext> |
<pad> | Padding within batches | Optional for pretraining, common for fine-tuning/inference batching |
<unk> | Unknown token (fallback) | Nearly unused in byte-level BPE |
<sep> / <cls> | Sentence boundary / classification aggregation point | BERT-style |
| `````` etc. | Dialogue template / role markers (ChatML style) | Qwen, many chat models |
The semantics of special tokens are defined by how they're used during training: for example, where <eos> appears and how dialogue role markers are organized must remain consistent across pretraining: data & objectives and fine-tuning: SFT and parameter-efficient fine-tuning; otherwise the model gets confused at inference time. This is a classic pitfall: "the tokenizer and training data protocol must be aligned."
An intuitive example of a ChatML-style dialogue template: <|im_start|>system\nYou are an assistant\n<|im_start|>user\nHello\n<|im_start|>assistant\n. These `````` markers should appear in the same form during pretraining data so the model learns to recognize "seeing this marker means a role switch"; if new special tokens are introduced only at fine-tuning time, the model has near-zero understanding of their semantics, requiring an additional adaptation period.
5. Vocabulary Size: How to Choose
Vocabulary size is the tokenizer's primary hyperparameter. Mainstream models provide these empirical values:
| Model | Vocab Size | Tokenization | Note |
|---|---|---|---|
| GPT-2 | 50,257 | byte-level BPE | 50k merges + 256 bytes + 1 special token |
| LLaMA / Llama 2 | 32,000 | SentencePiece (BPE) | Early Llama family |
| Llama 3 | 128,256 | tiktoken-style BPE | Significantly larger, multilingual improvement |
| GPT-4 (cl100k_base) | ~100,256 | byte-level BPE | ~100k range |
| Qwen2/Qwen2.5 | ~152,000 | BPE | China-centric vocabulary expansion (dataAsOf 2025, subject to official release) |
The selection logic is a triangular tradeoff:
- Large vocab → shorter sequences (saves attention), higher information density, but embedding/output layer parameters explode (parameters ≈ vocab size × hidden dim), and more data is needed to train sparse items adequately.
- Small vocab → saves parameters, but sequences get longer and rare words are chopped into fragments.
- Language structure → character-dense languages like Chinese and Japanese need larger vocabularies or finer subword splits; otherwise a single Hanzi character might be split into two or three tokens, increasing cost and losing semantics.
Another often-overlooked effect: vocabulary size is tied to "information per token": doubling the vocab doesn't necessarily halve the token count of text — it only reduces splitting granularity for the most common portion, while the long tail is still split the old way. In practice, the more common approach is: determine vocab scale based on the target language's character coverage needs, then confirm with a small proxy experiment (comparing average token counts across different vocab sizes on the same development set).
Vocabulary size affects "parameters per token" — not "one token = one parameter"
Output layer and embedding often share weights (tied embeddings). Expanding vocab from 32k to 128k adds roughly 96k × hidden dim parameters to the output head alone. When expanding vocab, the tokenizer should typically be retrained and vocab "aligned" — you can't just concatenate.
6. Tokenization Effects on Multilingual, Code, and Numbers
Tokenization's implicit impact is often underestimated in engineering:
| Scenario | Problem | Impact |
|---|---|---|
| Chinese/Japanese | No-space tokenization; BPE natively splits by character | Synonymous words spread across multiple tokens, making semantic modeling harder; vocab must be specifically expanded |
| Code | Complex indentation/symbols/identifier forms | Tokenizers not trained on code chop common code sequences into fragments |
| Numbers | "1234567" can be split into arbitrary chunks | Models' arithmetic and sorting abilities are affected by token splitting (digit misalignment) |
| Rare Proper Nouns | Names, places, rare characters | Semantic loss when split into fragments, especially noticeable for multilingual names |
| Noisy Text | Social media spelling variants | Each variant becomes an independent token, diluting statistics |
For example: the Chinese phrase "机器学习" (machine learning) might be 机器/学习 (two tokens) in a BPE with good coverage, but could become 机/器/学/习 (four tokens) or even include byte fragments in a poorly-covered tokenizer. Meanwhile, the same 4 characters in an English-centric model often only consume the token-equivalent of 2–3 English tokens. The direct consequence of doubled token count: with the same 128K context window, the Chinese content it can hold is roughly half that of English — cost and effective length both shrink.
A worth-remembering rule of thumb
The tokenizer's training corpus should come from the same source as the pretraining corpus: if pretraining data is 40% Chinese, the tokenizer training data should have a similar Chinese ratio. Otherwise Chinese gets chopped into inefficient long sequences, wasting cost. Many open-source model repos publish their tokenizer training corpus composition — worth reading.
7. Token Count, Cost, and Context Length Conversions
Tokens are the "currency" of billing and memory:
- Rough conversions: English ≈ 1 token per 0.75 words (1 word ≈ 1.3 tokens); Chinese ≈ 1 Hanzi character ≈ 1–2 tokens (depending on vocab); code-heavy scenes have higher token counts.
- Cost formula: total tokens per call = input tokens + output tokens. Most APIs bill on this basis (output often costs more per token).
- Context usage: model context length is measured in tokens (see Context & Long Context); KV Cache memory also scales linearly with input token count (see Inference Fundamentals: Autoregression and Sampling).
- Speed impact: per-token latency is roughly constant; longer sequences mean longer prefixes to process before generation, increasing first-token latency.
text
Cost estimation example (assuming $1/M tokens input, $3/M tokens output):
A 2000-character Chinese document ≈ 2000–4000 tokens
A response ≈ 500 tokens
Single call ≈ (3000×1 + 500×3) / 10^6 ≈ $0.0045
→ Estimate tokens before batch conversations; don't let "invisible per-unit billing" surprise your billTry it yourself with code (using OpenAI's tiktoken as an example):
python
import tiktoken
enc = tiktoken.get_encoding("cl100k_base") # encoding used by GPT-4 family
ids = enc.encode("机器学习与大模型")
print(len(ids), ids) # observe how many tokens this Chinese text is split intoExperiment with different languages, code snippets, and number strings to build an intuitive sense that "token counts vary wildly with language and format."
| Input | Approximate token count (cl100k_base scale) | Note |
|---|---|---|
| "Hello, world!" | ~4 | English short phrases approximate word-level tokenization |
| "你好,世界!" | ~6–8 | Hanzi typically 1–2 tokens/character |
| "Machine learning is great" | ~6 | Primarily word-level |
| A snippet of indented Python code | Higher than equivalent-length English | Indentation/symbols/identifiers get fragmented |
(Above are order-of-magnitude examples; actual values depend on the specific tokenizer.)
8. Trade-offs and Boundaries
- Tokenization is an "invisible but expensive" component: changing the tokenizer = retraining the entire model. Vocabulary, approach, and special tokens must be locked in before pretraining: data and objectives starts.
- Non-portable across models: different models use entirely different tokenizers; the same text's token count and cost aren't interchangeable.
- Tokenization approaches keep evolving: byte-level BPE, multilingual vocabulary expansion, domain-specific optimization for code/math (e.g., adding "digit tokens") are all active directions. More radical approaches (like MegaByte, byte-level language models) advocate skipping tokenization and modeling byte sequences directly, decoupling the strong coupling of "tokenizer and model must be retrained together" — currently uneconomical due to excessive sequence lengths, but it reminds us: tokenization isn't natural law; it's the optimal compromise under current compute constraints.
- Debugging entry point: if a model performs abnormally on a certain language or domain, the first step is always to inspect what the tokenizer splits the input into — this often directly exposes the root cause.
Practical advice
When debugging any abnormal model behavior, ask three questions first: what tokens does this text split into? what are each token's IDs? has any critical fragment been unexpectedly chopped? Tokenizer tools (Hugging Face tokenizers, tiktoken) give answers in minutes.
Tokenizer Design Decision Checklist
| Decision | Primary Considerations | Common Default |
|---|---|---|
| Algorithm | Coverage vs reversibility vs implementation cost | byte-level BPE / SentencePiece-Unigram |
| Vocab size | Parameter count, sequence length, language coverage | 32k–128k (larger for multilingual) |
| Byte-level? | Full Unicode coverage vs semantic granularity | Yes (no OOV) |
| Special token design | Alignment with training/dialogue protocol | ChatML or as-needed |
| Training corpus mix | Same source and ratio as pretraining corpus | Match primary corpus |
| Evaluation criteria | PPL and cost both depend on token counts | Use the same tokenizer for comparison |
How to Evaluate Tokenization Quality
Beyond efficiency metrics like "average token count," mature teams also check four things: roundtrip consistency (can encode→decode restore the original text losslessly?), multilingual coverage (average split length and OOV rate for target languages), special token completeness (one-to-one alignment with templates/protocols), and long-tail stability (consistent tokenization of the same word across different contexts). These checks can be scripted into a tokenizer-change CI pipeline to prevent "changed the vocab and the model quietly got worse."
Further Reading
- Language Modeling: The Next-Token Prediction Paradigm — tokens are the true unit of "the next word"
- Pretraining: Data and Objectives — the tokenizer alongside the corpus determines pretraining quality
- Context & Long Context — context length and cost are measured in tokens
- Inference Fundamentals: Autoregression and Sampling — how the token stream is generated step by step
- What Are Large Language Models? — the model's overall input-output pipeline
References
- Sennrich, Haddow, Birch. Neural Machine Translation of Rare Words with Subword Units (2016) — the original BPE tokenization paper
- Wu et al. Google's Neural Machine Translation System (2016, WordPiece) — the origin of WordPiece
- Kudo. Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates (Unigram, 2018) — the Unigram tokenization algorithm
- Kudo & Richardson. SentencePiece: A simple and language independent subword tokenizer (2018) — the SentencePiece framework paper
- Radford et al. Language Models are Unsupervised Multitask Learners (GPT-2, 2019) — introduction of byte-level BPE
- Hugging Face Tokenizers Documentation — high-performance tokenization implementation and tutorials
- SentencePiece GitHub — open-source implementation