Skip to content

Tokenization & Vocabulary

At a glance Tokenization chops continuous text into the smallest unit a model processes — tokens — and defines the true "input language" for language modeling. This article systematically compares character/word/BPE/WordPiece/Unigram/SentencePiece approaches, walks through the BPE algorithm step by step, and explains vocabulary size, special tokens, multilingual/code/number effects, and token-to-cost conversions.

Tokenization & Vocabulary ​

Tokenization is the process of chopping continuous text into discrete units — tokens — which are the smallest units a language model truly "sees" and processes. Models don't understand text directly; they only recognize integer IDs from their vocabulary (vocabulary). What language modeling calls "predicting the next word" is precisely "predicting the next token." Tokenization quality determines what the model can learn, how well it learns, and how much compute and cost it incurs — hence it's often called "the most underrated critical component."

One-line summary: the tokenizer defines the model's input alphabet — it determines "what the model can express" (vocabulary coverage) and "how much each sentence costs" (token count).

1. Why Tokenization Is Needed ​

A model could theoretically model at the character level, but pure character granularity has three problems:

ProblemExplanation
Sequences get too longAn average word has 4–5 characters, inflating sequence length several-fold and making self-attention O(n²) costs explode
Semantic units are fragmentedThe 5 characters of "apple" are independent; the model must learn to stitch them together, increasing learning burden
Knowledge doesn't cross tasksSubword-level knowledge (e.g., the grammatical meaning of "ing", "tion") cannot be directly reused by a character-level model

Tokenization seeks a balance between granularity and coverage: smaller units can cover any text (including rare words) but produce longer sequences with fragmented semantics; larger units preserve semantic integrity but lead to sparser vocabularies and harder new-word coverage. Modern mainstream approaches use subword tokenization: common words as wholes, rare words decomposed into composable subword fragments — the best of both worlds.

2. Tokenization Approaches Compared ​

ApproachBasic UnitVocabulary SourceProsConsTypical Use
Character (char)Individual charactersFixedTiny vocab, covers everythingLong sequences, no semanticsFallback for rare languages
WordFull wordsCorpus statisticsComplete semanticsHuge vocab, many OOVs, morphological explosionTraditional NLP
BPESubword (byte merging)Corpus-statistic mergingGood coverage, reversible, simple to implementNo linguistic priors, unstable for rare morphologiesGPT, Llama family
WordPieceSubword (probability merging)Corpus-statistic mergingBetter merge criterion (likelihood gain)Depends on space-pre-tokenized inputBERT, T5 family
UnigramSubword (probability reduction)Reverse-inferred from large vocabControllable vocab size, output can sample by probabilityComplex implementationSentencePiece, LLaMA (later)
SentencePieceSubword + whole-sentence encodingUnified framework of above methodsNo space pre-tokenization, byte-level support, cross-lingualMany concepts, requires configurationLLaMA, Mistral, Qwen family

Two core algorithms — BPE (Byte-Pair Encoding) and WordPiece — differ in their merge criteria: BPE always merges the most frequently co-occurring adjacent pair; WordPiece always merges the adjacent pair that maximizes the language model likelihood gain (like "best single-step tokenization"). Unigram goes in the opposite direction: first build an oversized candidate vocabulary, then iteratively remove words based on EM-computed loss contribution until reaching the target size. All three are based on the insight that "subwords = the solution to balancing vocab coverage and sequence length."

Notably, word-level tokenization hasn't been phased out: for morphologically simple languages (like English with regular word forms), it remains intuitive. But for morphologically rich languages (German compound words, Turkish agglutinative), spaceless languages (Chinese/Japanese/Thai), and high-OOV scenarios, subword tokenization's systematic advantages make it the de facto standard for modern LLMs.

3. BPE Algorithm: Step-by-Step Breakdown ​

BPE was introduced by Sennrich et al. 2016 from the data compression domain into neural machine translation, and is the mainstream algorithm adopted by the GPT series (byte-level variant: byte-level BPE) and the Llama family. The core idea: repeatedly merge the most frequently co-occurring adjacent token pair in the corpus, until the vocabulary reaches a target size.

text
Input: training corpus, target vocabulary size V
Step 0: Split each word into character sequences and count, e.g.
        "low"×5  "lower"×2  "newest"×6  "widest"×3
Step 1: Count all adjacent character pairs' frequencies:
        ("l","o")=7  ("o","w")=7  ("n","e")=6  ("w","e")=8  ...
Step 2: Merge the highest-frequency pair ("w","e") → new vocab entry "we"
        Corpus transforms: lo-w·er → "lo" "we" "r" form (illustrative)
Step 3: Repeat: count new adjacent pair frequencies, continue merging highest
        ("lo","w")=7 → "low"; ("n","e")=6 → "ne"; ("ne","w")=6 → "new"...
Step 4: Stop when vocab size reaches V, training complete

Encoding new text:
  "lowest" → greedy match longest vocab entries:
  Start at character level, repeatedly apply learned merge rules:
  l-o-w-e-s-t → "low" "est" (if "est" is also in vocab)

Key properties:

  • Byte-level BPE: starting from GPT-2, input is reduced to UTF-8 bytes before merging, so the vocabulary only depends on 256 byte values. No Unicode text will produce out-of-vocabulary (OOV) tokens — this is the foundation of multilingual robustness.
  • Deterministic greedy encoding: encoding applies greedy merges by longest match/learned rules, fast and reproducible.
  • Deficiency: rules are independent of language structure — whether morphological fragments like "st" and "tion" form, and in what order, depends entirely on frequency.

A complete example clarifies how byte-level BPE "grows": if the corpus heavily features "low" and "lower", training would first merge "lo" (if it's the highest-frequency pair), then form "low", and finally may form "lower"; while in "newest" and "widest", the high-frequency "est" is merged independently and reused for all words ending in -est. After training, "lowest" would be greedily split into two subwords: "low" + "est" — this is exactly how subword approaches work in practice: common words as wholes, rare words in chunks.

python
# Merge logic for byte-level BPE (pseudocode, illustrating the training loop)
def train_bpe(corpus: dict[str, int], vocab_size: int):
    vocab = set()                      # vocabulary
    while len(vocab) < vocab_size:
        pairs = count_adjacent_pairs(corpus)   # count adjacent token pair frequencies
        best = max(pairs, key=pairs.get)       # pick the highest-frequency pair
        vocab.add(best)                        # add to vocabulary
        corpus = merge_pair(corpus, best)      # replace that pair in corpus
    return vocab

Practical implementation

No need to implement from scratch in production — SentencePiece (Google, open source) unifies BPE/Unigram/WordPiece and eliminates space pre-tokenization; the Hugging Face tokenizers library provides a high-performance Rust implementation. Writing BPE from scratch is about understanding principles, not reinventing the wheel.

4. Special Tokens & Vocabulary Design ​

The vocabulary isn't just content tokens; it also contains special tokens used by the model's protocol:

Special TokenPurposeCommon Notation
<bos>Sequence startGPT-2 doesn't use; Llama does
<eos>Sequence end / dialogue turn boundaryGPT-style uses <endoftext>
<pad>Padding within batchesOptional for pretraining, common for fine-tuning/inference batching
<unk>Unknown token (fallback)Nearly unused in byte-level BPE
<sep> / <cls>Sentence boundary / classification aggregation pointBERT-style
`````` etc.Dialogue template / role markers (ChatML style)Qwen, many chat models

The semantics of special tokens are defined by how they're used during training: for example, where <eos> appears and how dialogue role markers are organized must remain consistent across pretraining: data & objectives and fine-tuning: SFT and parameter-efficient fine-tuning; otherwise the model gets confused at inference time. This is a classic pitfall: "the tokenizer and training data protocol must be aligned."

An intuitive example of a ChatML-style dialogue template: <|im_start|>system\nYou are an assistant\n<|im_start|>user\nHello\n<|im_start|>assistant\n. These `````` markers should appear in the same form during pretraining data so the model learns to recognize "seeing this marker means a role switch"; if new special tokens are introduced only at fine-tuning time, the model has near-zero understanding of their semantics, requiring an additional adaptation period.

5. Vocabulary Size: How to Choose ​

Vocabulary size is the tokenizer's primary hyperparameter. Mainstream models provide these empirical values:

ModelVocab SizeTokenizationNote
GPT-250,257byte-level BPE50k merges + 256 bytes + 1 special token
LLaMA / Llama 232,000SentencePiece (BPE)Early Llama family
Llama 3128,256tiktoken-style BPESignificantly larger, multilingual improvement
GPT-4 (cl100k_base)~100,256byte-level BPE~100k range
Qwen2/Qwen2.5~152,000BPEChina-centric vocabulary expansion (dataAsOf 2025, subject to official release)

The selection logic is a triangular tradeoff:

  • Large vocab → shorter sequences (saves attention), higher information density, but embedding/output layer parameters explode (parameters ≈ vocab size × hidden dim), and more data is needed to train sparse items adequately.
  • Small vocab → saves parameters, but sequences get longer and rare words are chopped into fragments.
  • Language structure → character-dense languages like Chinese and Japanese need larger vocabularies or finer subword splits; otherwise a single Hanzi character might be split into two or three tokens, increasing cost and losing semantics.

Another often-overlooked effect: vocabulary size is tied to "information per token": doubling the vocab doesn't necessarily halve the token count of text — it only reduces splitting granularity for the most common portion, while the long tail is still split the old way. In practice, the more common approach is: determine vocab scale based on the target language's character coverage needs, then confirm with a small proxy experiment (comparing average token counts across different vocab sizes on the same development set).

Vocabulary size affects "parameters per token" — not "one token = one parameter"

Output layer and embedding often share weights (tied embeddings). Expanding vocab from 32k to 128k adds roughly 96k × hidden dim parameters to the output head alone. When expanding vocab, the tokenizer should typically be retrained and vocab "aligned" — you can't just concatenate.

6. Tokenization Effects on Multilingual, Code, and Numbers ​

Tokenization's implicit impact is often underestimated in engineering:

ScenarioProblemImpact
Chinese/JapaneseNo-space tokenization; BPE natively splits by characterSynonymous words spread across multiple tokens, making semantic modeling harder; vocab must be specifically expanded
CodeComplex indentation/symbols/identifier formsTokenizers not trained on code chop common code sequences into fragments
Numbers"1234567" can be split into arbitrary chunksModels' arithmetic and sorting abilities are affected by token splitting (digit misalignment)
Rare Proper NounsNames, places, rare charactersSemantic loss when split into fragments, especially noticeable for multilingual names
Noisy TextSocial media spelling variantsEach variant becomes an independent token, diluting statistics

For example: the Chinese phrase "机器学习" (machine learning) might be 机器/学习 (two tokens) in a BPE with good coverage, but could become 机/器/学/习 (four tokens) or even include byte fragments in a poorly-covered tokenizer. Meanwhile, the same 4 characters in an English-centric model often only consume the token-equivalent of 2–3 English tokens. The direct consequence of doubled token count: with the same 128K context window, the Chinese content it can hold is roughly half that of English — cost and effective length both shrink.

A worth-remembering rule of thumb

The tokenizer's training corpus should come from the same source as the pretraining corpus: if pretraining data is 40% Chinese, the tokenizer training data should have a similar Chinese ratio. Otherwise Chinese gets chopped into inefficient long sequences, wasting cost. Many open-source model repos publish their tokenizer training corpus composition — worth reading.

7. Token Count, Cost, and Context Length Conversions ​

Tokens are the "currency" of billing and memory:

  • Rough conversions: English ≈ 1 token per 0.75 words (1 word ≈ 1.3 tokens); Chinese ≈ 1 Hanzi character ≈ 1–2 tokens (depending on vocab); code-heavy scenes have higher token counts.
  • Cost formula: total tokens per call = input tokens + output tokens. Most APIs bill on this basis (output often costs more per token).
  • Context usage: model context length is measured in tokens (see Context & Long Context); KV Cache memory also scales linearly with input token count (see Inference Fundamentals: Autoregression and Sampling).
  • Speed impact: per-token latency is roughly constant; longer sequences mean longer prefixes to process before generation, increasing first-token latency.
text
Cost estimation example (assuming $1/M tokens input, $3/M tokens output):
  A 2000-character Chinese document ≈ 2000–4000 tokens
  A response ≈ 500 tokens
  Single call ≈ (3000×1 + 500×3) / 10^6 ≈ $0.0045
  → Estimate tokens before batch conversations; don't let "invisible per-unit billing" surprise your bill

Try it yourself with code (using OpenAI's tiktoken as an example):

python
import tiktoken
enc = tiktoken.get_encoding("cl100k_base")   # encoding used by GPT-4 family
ids = enc.encode("机器学习与大模型")
print(len(ids), ids)   # observe how many tokens this Chinese text is split into

Experiment with different languages, code snippets, and number strings to build an intuitive sense that "token counts vary wildly with language and format."

InputApproximate token count (cl100k_base scale)Note
"Hello, world!"~4English short phrases approximate word-level tokenization
"你好,世界!"~6–8Hanzi typically 1–2 tokens/character
"Machine learning is great"~6Primarily word-level
A snippet of indented Python codeHigher than equivalent-length EnglishIndentation/symbols/identifiers get fragmented

(Above are order-of-magnitude examples; actual values depend on the specific tokenizer.)

8. Trade-offs and Boundaries ​

  • Tokenization is an "invisible but expensive" component: changing the tokenizer = retraining the entire model. Vocabulary, approach, and special tokens must be locked in before pretraining: data and objectives starts.
  • Non-portable across models: different models use entirely different tokenizers; the same text's token count and cost aren't interchangeable.
  • Tokenization approaches keep evolving: byte-level BPE, multilingual vocabulary expansion, domain-specific optimization for code/math (e.g., adding "digit tokens") are all active directions. More radical approaches (like MegaByte, byte-level language models) advocate skipping tokenization and modeling byte sequences directly, decoupling the strong coupling of "tokenizer and model must be retrained together" — currently uneconomical due to excessive sequence lengths, but it reminds us: tokenization isn't natural law; it's the optimal compromise under current compute constraints.
  • Debugging entry point: if a model performs abnormally on a certain language or domain, the first step is always to inspect what the tokenizer splits the input into — this often directly exposes the root cause.

Practical advice

When debugging any abnormal model behavior, ask three questions first: what tokens does this text split into? what are each token's IDs? has any critical fragment been unexpectedly chopped? Tokenizer tools (Hugging Face tokenizers, tiktoken) give answers in minutes.

Tokenizer Design Decision Checklist ​

DecisionPrimary ConsiderationsCommon Default
AlgorithmCoverage vs reversibility vs implementation costbyte-level BPE / SentencePiece-Unigram
Vocab sizeParameter count, sequence length, language coverage32k–128k (larger for multilingual)
Byte-level?Full Unicode coverage vs semantic granularityYes (no OOV)
Special token designAlignment with training/dialogue protocolChatML or as-needed
Training corpus mixSame source and ratio as pretraining corpusMatch primary corpus
Evaluation criteriaPPL and cost both depend on token countsUse the same tokenizer for comparison

How to Evaluate Tokenization Quality ​

Beyond efficiency metrics like "average token count," mature teams also check four things: roundtrip consistency (can encode→decode restore the original text losslessly?), multilingual coverage (average split length and OOV rate for target languages), special token completeness (one-to-one alignment with templates/protocols), and long-tail stability (consistent tokenization of the same word across different contexts). These checks can be scripted into a tokenizer-change CI pipeline to prevent "changed the vocab and the model quietly got worse."

Further Reading ​

References ​