Appearance
Distillation Classics: KD and Its Family
If quantization compresses a model, distillation (knowledge distillation) copies the knowledge inside a large model into a small one — the former changes numerical precision, the latter changes model capacity. The philosophy of distillation: a model's size is not the measure of its knowledge; knowledge can be transferred. This page walks through the seminal Hinton KD paper, then breaks down the three BERT-family siblings one by one (DistilBERT, TinyBERT, MiniLM).
Read Model Compression first to position distillation within the engineering landscape, then come back for the details.
Distillation Overview: Why Soft Labels Carry More Information Than Hard Labels
After a classification model is trained, it outputs a probability distribution for input x. Traditional training computes cross-entropy against hard labels (one-hot, like [0,0,1]); distillation instead uses the teacher's soft outputs as the supervision signal. What's the difference?
Take "is this image a cat?": a hard label only tells the model "cat." But a trained teacher outputs something like [cat: 0.7, dog: 0.2, rabbit: 0.1] — implicitly encoding the inter-class similarity structure the teacher learned from the data, such as "cats resemble dogs far more than they resemble cars." Hard labels throw this soft information away entirely. Distillation = picking that discarded information back up and feeding it to the student.
Temperature T is distillation's key knob. The teacher's logits z are softened before being handed to the student:
text
Hard label: [0, 0, 1] (one-hot, least information)
teacher output p_i = exp(z_i / T) / Σ_j exp(z_j / T)
T = 1 : original softmax — the distribution already shows relative magnitudes
T → large: the distribution flattens, amplifying tiny inter-class relationships
T → small: the distribution sharpens toward one-hot
Train with a larger T (e.g. 2–8) so the student learns the structure;
at inference T=1 (the student outputs normally) — temperature belongs to training only.The student's total loss = distillation loss (student soft output vs. teacher soft output, at the same temperature) + the student's own hard-label cross-entropy, combined as a weighted sum. The hard-label term keeps the student from absorbing the teacher's biases.
1. Hinton KD: The Founding Paper (Hinton, Vinyals & Dean, 2015)
The contribution in one sentence
Knowledge transfer via temperature-softened soft labels: a small model (the student) trained under joint supervision from soft targets plus hard labels converges better than direct training — and knowledge distillation became a research field of its own.
Background and motivation
Around 2015, model compression meant pruning and quantization — both "subtraction on an existing model." Hinton proposed a more fundamental idea: training a large model (or an ensemble) costs enormous compute, but can its knowledge be transplanted directly into a small model? The paper validated the idea on three tasks — MNIST digits, speech recognition (acoustic models), and machine translation (WMT data) — with students matching or approaching their teachers.
The method in brief
- Soft labels as supervision: the teacher's outputs are softened with temperature T, and the student computes cross-entropy against the softened distribution.
- Choosing T: typically the teacher itself is first trained with a larger T (making its soft labels more informative), then the student is trained at the same T; at inference, back to T=1.
- Ensemble distillation: use an ensemble of multiple models as the teacher and distill into a single student — one student model captures the ensemble's accuracy.
text
┌─────────────────────────┐
│ Teacher (large model / │
│ ensemble) │
│ outputs logits z │
└───────────┬─────────────┘
│ softmax(z / T)
▼
soft labels ┌── student hard-label CE (bias guard)
p_teacher │
│ ▼
▼
student soft output p_student ──CE── student's own training data
│
┌──────────┴──────────┐
│ distill loss + task │
│ loss │
└─────────────────────┘Key results
- MNIST: an ensemble of 8 models distilled into one — error rate on par with the ensemble at 1/8 the inference cost.
- Speech recognition (acoustic models): the distilled small model approximated ensemble accuracy.
- The enduring insight: "model size" and "knowledge capacity" are two separate variables — soft labels let the student encode the same knowledge with fewer parameters.
Limitations and follow-ups
- Validated mainly on classification and sequence models; distilling pretrained Transformers required new distillation objectives (the siblings below).
- Effectiveness depends on teacher quality and the temperature hyperparameter.
- Direct descendants: fitnets (distilling intermediate layers), two-stage teacher-assistant distillation, and the BERT-family distillations.
What it means today
KD defined the paradigm of trading training cost for inference cost: as long as you can afford to train a large model (or simply use an existing one for offline labeling), you can ship a much cheaper small model. Today's LLM practice of "using a large model to generate data and train a small one" (Alpaca, Self-Instruct) is essentially Hinton KD's idea reborn for generative tasks.
2. DistilBERT: Distilling at Pretraining Time (Sanh et al., 2019)
The contribution in one sentence
Moving distillation upstream into pretraining: a small BERT is taught by a large BERT's soft outputs during pretraining itself, producing a general-purpose model that is 40% smaller, 60% faster, and retains 97% of the performance — and fine-tunes as usual.
Background and motivation
Hinton KD was commonly applied as task-level distillation: train a task model, then distill it into a small one. But BERT follows the "pretrain + fine-tune" paradigm — distilling only at fine-tuning time means the small model still has to acquire its language knowledge through its own (expensive) pretraining. DistilBERT's move: distill during pretraining, producing a general-purpose model that is "born small," then fine-tune downstream as usual.
The method in brief
- Structure: BERT-base's 12 layers → a 6-layer student (initialized by taking every other layer), ~66M parameters (down from 110M).
- Triple loss: a language modeling loss (the student's own masked LM) + distillation loss (cross-entropy between student/teacher softmax outputs at temperature T=2) + cosine distance loss (aligning the direction of student/teacher hidden states).
- One loss simultaneously handles "learn the task + learn the teacher's soft targets + learn the teacher's representation direction."
Key results
| Metric | Number |
|---|---|
| Parameters | 40% fewer (110M → 66M, 6 layers) |
| Performance | Retains 97% of language understanding (GLUE multi-task average) |
| Speed | Inference ~60% faster |
| Applicability | After fine-tuning, covers most NLP tasks; usable directly as a general model |
Limitations and follow-ups
- Only output-level soft targets plus one hidden-state alignment — attention matrices are not distilled; later work argued attention structure information is just as critical.
- 6 layers is the typical floor; deeper compression needs finer distillation objectives.
What it means today
DistilBERT proved that pretraining-time distillation is enormously valuable in practice: distill once into a general-purpose small model, and every downstream fine-tuning job reuses it — the inference savings compound indefinitely. In BERT deployment, distilbert-* remains one of the default "fast + reliable" choices to this day.
3. TinyBERT: Two Stages Plus Attention-Matrix Distillation (Jiao et al., EMNLP 2020 Findings)
The contribution in one sentence
Refining distillation into four levels of objectives — embedding layer, attention matrices, hidden states, and prediction layer — executed in two stages (pretraining and fine-tuning): a 4-layer TinyBERT reaches 96.8% of BERT-base's performance while being 7.5× smaller and 9.4× faster.
Background and motivation
DistilBERT showed output-level alignment works, but compressing to 4 layers (a third of the depth) demands more. The paper's argument: much of a Transformer's capability lives in its attention structure (which tokens each token looks at, and how heavily) — knowledge that is only implicit in soft labels. TinyBERT takes the attention matrix (the QK^T distribution) itself as a distillation objective, making the student's attention "look like" the teacher's.
The method in brief
- Four levels of distillation objectives:
- Embedding-layer output alignment;
- Attention matrix alignment (each student layer's attention distribution ≈ the teacher's corresponding layer — the core innovation);
- Hidden state alignment (including linearly transformed student/teacher hidden states);
- Output alignment (soft labels).
- A two-stage framework: first distill once during generic-corpus pretraining (yielding a general small model), then distill again per downstream task during task fine-tuning — "generic distillation as the base, task distillation for the polish."
- Layer mapping: the 4-layer student maps onto the teacher's 12 layers at intervals (e.g. 1→[1,2,3], 2→[4,5,6]...).
Key results
| Metric | Number |
|---|---|
| Size | 4-layer TinyBERT (~14.5M) vs. BERT-base (110M): ~7.5× smaller |
| Performance | 96.8% of the teacher on GLUE; the 6-layer version matches BERT-base |
| Speed | Inference ~9.4× faster |
| Comparison | The 4-layer version clearly beat other 4-layer distilled models of the time (with ~28% of their parameters) |
Limitations and follow-ups
- Two-stage distillation = higher training cost (the pretraining stage runs distillation too).
- Attention alignment adds hyperparameters (layer-mapping scheme, alignment weights).
- Follow-ups: MiniLM observed that layer-by-layer attention alignment suffices for 12→6 but has room to improve for deeper compression (12→4), and proposed the simpler "align only the last layer."
What it means today
TinyBERT is the exemplar of "distilling the model's internal structure": when compressing a model to a third of its depth, don't just align outputs — align attention and representations too. Virtually every later Transformer distillation (including LLM distillation) inherited this idea.
4. MiniLM: Distilling Only the Self-Attention's Internal Relations (Wang et al., EMNLP 2020 Findings)
The contribution in one sentence
The distillation objective is abstracted further to the self-attention's "internal relations" (QK^T distributions plus V^T V similarities), aligned only against the teacher's last layer — depth scales freely (12→6 or 12→4), and 50% of the parameters retain 99%+ accuracy.
Background and motivation
TinyBERT's layer-by-layer attention alignment has a problem: the student's depth/width differ from the teacher's, so the layer mapping is heuristic and every layer needs its own alignment design. MiniLM proposed a more "depth-agnostic" scheme: take only the teacher's last-layer self-attention module as the single distillation target, and have every student layer align to it.
More importantly, it distills not just the attention distribution (QK^T) but adds V^T V — the relations among values in self-attention, which carry the structural information of "which token representations should be stacked together."
The method in brief
text
Teacher (last layer) Student (any depth)
│ │
self-attention self-attention
module module
│ │
QK^T (attention dist.) ←── align ──→ QK^T (attention dist.)
V^T V (value relations) ←── align ──→ V^T V (value relations)- Depth scaling: a 12-layer teacher can directly teach a 6- or 4-layer student, no layer mapping needed.
- Teacher Assistant: when the teacher is very large (e.g. 24 layers), first distill an "assistant," then distill the assistant into the student, narrowing the capacity gap.
Key results
| Metric | Number |
|---|---|
| Size | 12-layer BERT-base → 6 layers (50% of parameters) |
| Performance | 99%+ accuracy retained across SQuAD 2.0 and GLUE tasks |
| Compression axes | Supports depth scaling (12→6, 12→4) and width scaling; extended to multilingual models |
| Versus peers | Outperformed DistilBERT, TinyBERT, and other baselines at equal student size |
Limitations and follow-ups
- Only self-attention is distilled; MLP knowledge is covered indirectly through output soft labels — theoretically not the full picture.
- Teacher Assistant requires training an extra intermediate model.
- Follow-ups: MiniLM's ideas (relational distillation, depth-agnostic alignment) were widely inherited by LLM distillation (e.g. token-level distillation).
What it means today
MiniLM showed that a distillation objective isn't better for being more exhaustive — it's better for being more structural. Capturing "how the model organizes information" (attention relations) beats layer-by-layer, target-by-target alignment. The deployment lesson: when resources are tight, prioritize preserving the model's information organization, not bit-exact numerical alignment.
5. Side-by-Side Comparison
| Model | Parameters | Performance retained | Inference speedup | Distillation objectives | Distillation stage |
|---|---|---|---|---|---|
| DistilBERT | 40% (110M→66M) | ~97% | 60% faster | Soft labels + hidden-state cosine | Pretraining |
| TinyBERT (4-layer) | ~1/7.5 (vs. BERT-base) | 96.8% (6-layer matches) | 9.4× faster | Embedding + attention matrices + hidden states + output | Pretraining + task (two-stage) |
| MiniLM | 50% (12→6 layers) | 99%+ | ~2× | Last-layer self-attention relations (QK^T + V^T V) | Pretraining-style, depth-agnostic |
6. Distillation and Quantization: A One-Two Punch
Distillation and quantization are orthogonal and stack well: distill to a small model first, then quantize the small model. This usually yields better accuracy than "quantizing the large model directly" and lower memory than "distilling without quantizing."
- Why distill first: small models have tidier weight distributions with fewer outliers, making quantization error more controllable (compare LLM.int8's outlier problem).
- Engineering order: distill to cut capacity → quantize to cut bit width → graph optimization/operator fusion to cut latency.
- Caution: the quantization calibration set should come from the distilled model's target distribution — don't force the teacher's distribution onto it.
The deployment-side decision order
Tight on both latency and memory, model under 1B: distill first, quantize if still not enough. Very large model (≥7B) that can't be retrained: quantize directly (GPTQ/AWQ) — see Quantization Classics.
7. Implications for LLMs: Distill or Just Train Smaller?
Two questions that often get conflated:
- How is distillation different from "just training a small model"? Training small from scratch means learning all knowledge from zero; distillation lets the small model inherit knowledge already refined by the large model (soft labels + structural alignment), so at equal capacity it converges better and lands closer to the large model.
- Does distillation still matter in the LLM era? Yes, but the form has changed:
- Data distillation: use a large model offline to generate high-quality data/responses and train a small model (the Alpaca / Self-Instruct approach) — the mainstream route to small LLMs today;
- Online/offline generation distillation: teacher models guiding students to align on generation tasks (requiring new solutions for problems like distribution collapse);
- Distilling into stronger small models: MiniLM's teacher-assistant idea continues to pay off in multi-stage distillation.
- Distillation ≠ a replacement for quantization: they answer different questions — distillation cuts capacity, quantization cuts bit width. Capacity and bit width are independent dimensions.
8. Shared Limitations
- Training cost comes first: distillation is fundamentally "training cost for inference cost" — without the budget, the payoff shrinks.
- Capacity gaps: when the teacher-student gap is too large, the student can't keep up; teacher assistants or staged distillation are needed.
- Task dependence: task-level distillation may generalize poorly to new tasks; generic distillation (the DistilBERT/MiniLM route) is more robust but with milder gains.
- Beware "teaching to the test": on small test sets, a student can overfit the teacher's noise; evaluate with an independent test set and calibrated distributions.
Further Reading
- Model Compression — the full engineering picture of distillation, pruning, and low-rank methods
- Quantization Classics: LLM.int8 / GPTQ / AWQ — the distillation + quantization combo
- Paper Map — where the distillation family sits in the field
- Model Formats and Conversion — how distilled artifacts get exported and ported across frameworks
- Deploy a Model from Scratch — a hands-on route to shipping a small model end to end
References
- Distilling the Knowledge in a Neural Network (arXiv 1503.02531)
- DistilBERT, a distilled version of BERT (arXiv 1910.01108)
- TinyBERT: Distilling BERT for Natural Language Understanding (arXiv 1909.10351)
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers (arXiv 2002.10957)