Skip to content

Distillation Classics: KD and Its Family

At a glance Knowledge distillation has a large model teach a small one — the core technique for trading training cost for inference cost in deployment. This page walks through Hinton's original KD and the BERT-family distillations (DistilBERT, TinyBERT, MiniLM).

Distillation Classics: KD and Its Family ​

If quantization compresses a model, distillation (knowledge distillation) copies the knowledge inside a large model into a small one — the former changes numerical precision, the latter changes model capacity. The philosophy of distillation: a model's size is not the measure of its knowledge; knowledge can be transferred. This page walks through the seminal Hinton KD paper, then breaks down the three BERT-family siblings one by one (DistilBERT, TinyBERT, MiniLM).

Read Model Compression first to position distillation within the engineering landscape, then come back for the details.

Distillation Overview: Why Soft Labels Carry More Information Than Hard Labels ​

After a classification model is trained, it outputs a probability distribution for input x. Traditional training computes cross-entropy against hard labels (one-hot, like [0,0,1]); distillation instead uses the teacher's soft outputs as the supervision signal. What's the difference?

Take "is this image a cat?": a hard label only tells the model "cat." But a trained teacher outputs something like [cat: 0.7, dog: 0.2, rabbit: 0.1] — implicitly encoding the inter-class similarity structure the teacher learned from the data, such as "cats resemble dogs far more than they resemble cars." Hard labels throw this soft information away entirely. Distillation = picking that discarded information back up and feeding it to the student.

Temperature T is distillation's key knob. The teacher's logits z are softened before being handed to the student:

text
Hard label:  [0, 0, 1]                        (one-hot, least information)
teacher output p_i = exp(z_i / T) / Σ_j exp(z_j / T)

T = 1   : original softmax — the distribution already shows relative magnitudes
T → large: the distribution flattens, amplifying tiny inter-class relationships
T → small: the distribution sharpens toward one-hot

Train with a larger T (e.g. 2–8) so the student learns the structure;
at inference T=1 (the student outputs normally) — temperature belongs to training only.

The student's total loss = distillation loss (student soft output vs. teacher soft output, at the same temperature) + the student's own hard-label cross-entropy, combined as a weighted sum. The hard-label term keeps the student from absorbing the teacher's biases.

1. Hinton KD: The Founding Paper (Hinton, Vinyals & Dean, 2015) ​

The contribution in one sentence ​

Knowledge transfer via temperature-softened soft labels: a small model (the student) trained under joint supervision from soft targets plus hard labels converges better than direct training — and knowledge distillation became a research field of its own.

Background and motivation ​

Around 2015, model compression meant pruning and quantization — both "subtraction on an existing model." Hinton proposed a more fundamental idea: training a large model (or an ensemble) costs enormous compute, but can its knowledge be transplanted directly into a small model? The paper validated the idea on three tasks — MNIST digits, speech recognition (acoustic models), and machine translation (WMT data) — with students matching or approaching their teachers.

The method in brief ​

  • Soft labels as supervision: the teacher's outputs are softened with temperature T, and the student computes cross-entropy against the softened distribution.
  • Choosing T: typically the teacher itself is first trained with a larger T (making its soft labels more informative), then the student is trained at the same T; at inference, back to T=1.
  • Ensemble distillation: use an ensemble of multiple models as the teacher and distill into a single student — one student model captures the ensemble's accuracy.
text
         ┌─────────────────────────┐
         │  Teacher (large model /  │
         │  ensemble)               │
         │  outputs logits z        │
         └───────────┬─────────────┘
                     │ softmax(z / T)
                     ▼
                soft labels          ┌── student hard-label CE (bias guard)
                p_teacher            │
                     │               ▼
                     ▼
   student soft output p_student ──CE── student's own training data
                     │
          ┌──────────┴──────────┐
          │  distill loss + task │
          │       loss           │
          └─────────────────────┘

Key results ​

  • MNIST: an ensemble of 8 models distilled into one — error rate on par with the ensemble at 1/8 the inference cost.
  • Speech recognition (acoustic models): the distilled small model approximated ensemble accuracy.
  • The enduring insight: "model size" and "knowledge capacity" are two separate variables — soft labels let the student encode the same knowledge with fewer parameters.

Limitations and follow-ups ​

  • Validated mainly on classification and sequence models; distilling pretrained Transformers required new distillation objectives (the siblings below).
  • Effectiveness depends on teacher quality and the temperature hyperparameter.
  • Direct descendants: fitnets (distilling intermediate layers), two-stage teacher-assistant distillation, and the BERT-family distillations.

What it means today ​

KD defined the paradigm of trading training cost for inference cost: as long as you can afford to train a large model (or simply use an existing one for offline labeling), you can ship a much cheaper small model. Today's LLM practice of "using a large model to generate data and train a small one" (Alpaca, Self-Instruct) is essentially Hinton KD's idea reborn for generative tasks.

2. DistilBERT: Distilling at Pretraining Time (Sanh et al., 2019) ​

The contribution in one sentence ​

Moving distillation upstream into pretraining: a small BERT is taught by a large BERT's soft outputs during pretraining itself, producing a general-purpose model that is 40% smaller, 60% faster, and retains 97% of the performance — and fine-tunes as usual.

Background and motivation ​

Hinton KD was commonly applied as task-level distillation: train a task model, then distill it into a small one. But BERT follows the "pretrain + fine-tune" paradigm — distilling only at fine-tuning time means the small model still has to acquire its language knowledge through its own (expensive) pretraining. DistilBERT's move: distill during pretraining, producing a general-purpose model that is "born small," then fine-tune downstream as usual.

The method in brief ​

  • Structure: BERT-base's 12 layers → a 6-layer student (initialized by taking every other layer), ~66M parameters (down from 110M).
  • Triple loss: a language modeling loss (the student's own masked LM) + distillation loss (cross-entropy between student/teacher softmax outputs at temperature T=2) + cosine distance loss (aligning the direction of student/teacher hidden states).
  • One loss simultaneously handles "learn the task + learn the teacher's soft targets + learn the teacher's representation direction."

Key results ​

MetricNumber
Parameters40% fewer (110M → 66M, 6 layers)
PerformanceRetains 97% of language understanding (GLUE multi-task average)
SpeedInference ~60% faster
ApplicabilityAfter fine-tuning, covers most NLP tasks; usable directly as a general model

Limitations and follow-ups ​

  • Only output-level soft targets plus one hidden-state alignment — attention matrices are not distilled; later work argued attention structure information is just as critical.
  • 6 layers is the typical floor; deeper compression needs finer distillation objectives.

What it means today ​

DistilBERT proved that pretraining-time distillation is enormously valuable in practice: distill once into a general-purpose small model, and every downstream fine-tuning job reuses it — the inference savings compound indefinitely. In BERT deployment, distilbert-* remains one of the default "fast + reliable" choices to this day.

3. TinyBERT: Two Stages Plus Attention-Matrix Distillation (Jiao et al., EMNLP 2020 Findings) ​

The contribution in one sentence ​

Refining distillation into four levels of objectives — embedding layer, attention matrices, hidden states, and prediction layer — executed in two stages (pretraining and fine-tuning): a 4-layer TinyBERT reaches 96.8% of BERT-base's performance while being 7.5× smaller and 9.4× faster.

Background and motivation ​

DistilBERT showed output-level alignment works, but compressing to 4 layers (a third of the depth) demands more. The paper's argument: much of a Transformer's capability lives in its attention structure (which tokens each token looks at, and how heavily) — knowledge that is only implicit in soft labels. TinyBERT takes the attention matrix (the QK^T distribution) itself as a distillation objective, making the student's attention "look like" the teacher's.

The method in brief ​

  • Four levels of distillation objectives:
    1. Embedding-layer output alignment;
    2. Attention matrix alignment (each student layer's attention distribution ≈ the teacher's corresponding layer — the core innovation);
    3. Hidden state alignment (including linearly transformed student/teacher hidden states);
    4. Output alignment (soft labels).
  • A two-stage framework: first distill once during generic-corpus pretraining (yielding a general small model), then distill again per downstream task during task fine-tuning — "generic distillation as the base, task distillation for the polish."
  • Layer mapping: the 4-layer student maps onto the teacher's 12 layers at intervals (e.g. 1→[1,2,3], 2→[4,5,6]...).

Key results ​

MetricNumber
Size4-layer TinyBERT (~14.5M) vs. BERT-base (110M): ~7.5× smaller
Performance96.8% of the teacher on GLUE; the 6-layer version matches BERT-base
SpeedInference ~9.4× faster
ComparisonThe 4-layer version clearly beat other 4-layer distilled models of the time (with ~28% of their parameters)

Limitations and follow-ups ​

  • Two-stage distillation = higher training cost (the pretraining stage runs distillation too).
  • Attention alignment adds hyperparameters (layer-mapping scheme, alignment weights).
  • Follow-ups: MiniLM observed that layer-by-layer attention alignment suffices for 12→6 but has room to improve for deeper compression (12→4), and proposed the simpler "align only the last layer."

What it means today ​

TinyBERT is the exemplar of "distilling the model's internal structure": when compressing a model to a third of its depth, don't just align outputs — align attention and representations too. Virtually every later Transformer distillation (including LLM distillation) inherited this idea.

4. MiniLM: Distilling Only the Self-Attention's Internal Relations (Wang et al., EMNLP 2020 Findings) ​

The contribution in one sentence ​

The distillation objective is abstracted further to the self-attention's "internal relations" (QK^T distributions plus V^T V similarities), aligned only against the teacher's last layer — depth scales freely (12→6 or 12→4), and 50% of the parameters retain 99%+ accuracy.

Background and motivation ​

TinyBERT's layer-by-layer attention alignment has a problem: the student's depth/width differ from the teacher's, so the layer mapping is heuristic and every layer needs its own alignment design. MiniLM proposed a more "depth-agnostic" scheme: take only the teacher's last-layer self-attention module as the single distillation target, and have every student layer align to it.

More importantly, it distills not just the attention distribution (QK^T) but adds V^T V — the relations among values in self-attention, which carry the structural information of "which token representations should be stacked together."

The method in brief ​

text
Teacher (last layer)                    Student (any depth)
    │                                       │
 self-attention                         self-attention
    module                                 module
    │                                       │
 QK^T (attention dist.)  ←── align ──→  QK^T (attention dist.)
 V^T V (value relations) ←── align ──→  V^T V (value relations)
  • Depth scaling: a 12-layer teacher can directly teach a 6- or 4-layer student, no layer mapping needed.
  • Teacher Assistant: when the teacher is very large (e.g. 24 layers), first distill an "assistant," then distill the assistant into the student, narrowing the capacity gap.

Key results ​

MetricNumber
Size12-layer BERT-base → 6 layers (50% of parameters)
Performance99%+ accuracy retained across SQuAD 2.0 and GLUE tasks
Compression axesSupports depth scaling (12→6, 12→4) and width scaling; extended to multilingual models
Versus peersOutperformed DistilBERT, TinyBERT, and other baselines at equal student size

Limitations and follow-ups ​

  • Only self-attention is distilled; MLP knowledge is covered indirectly through output soft labels — theoretically not the full picture.
  • Teacher Assistant requires training an extra intermediate model.
  • Follow-ups: MiniLM's ideas (relational distillation, depth-agnostic alignment) were widely inherited by LLM distillation (e.g. token-level distillation).

What it means today ​

MiniLM showed that a distillation objective isn't better for being more exhaustive — it's better for being more structural. Capturing "how the model organizes information" (attention relations) beats layer-by-layer, target-by-target alignment. The deployment lesson: when resources are tight, prioritize preserving the model's information organization, not bit-exact numerical alignment.

5. Side-by-Side Comparison ​

ModelParametersPerformance retainedInference speedupDistillation objectivesDistillation stage
DistilBERT40% (110M→66M)~97%60% fasterSoft labels + hidden-state cosinePretraining
TinyBERT (4-layer)~1/7.5 (vs. BERT-base)96.8% (6-layer matches)9.4× fasterEmbedding + attention matrices + hidden states + outputPretraining + task (two-stage)
MiniLM50% (12→6 layers)99%+~2×Last-layer self-attention relations (QK^T + V^T V)Pretraining-style, depth-agnostic

6. Distillation and Quantization: A One-Two Punch ​

Distillation and quantization are orthogonal and stack well: distill to a small model first, then quantize the small model. This usually yields better accuracy than "quantizing the large model directly" and lower memory than "distilling without quantizing."

  • Why distill first: small models have tidier weight distributions with fewer outliers, making quantization error more controllable (compare LLM.int8's outlier problem).
  • Engineering order: distill to cut capacity → quantize to cut bit width → graph optimization/operator fusion to cut latency.
  • Caution: the quantization calibration set should come from the distilled model's target distribution — don't force the teacher's distribution onto it.

The deployment-side decision order

Tight on both latency and memory, model under 1B: distill first, quantize if still not enough. Very large model (≥7B) that can't be retrained: quantize directly (GPTQ/AWQ) — see Quantization Classics.

7. Implications for LLMs: Distill or Just Train Smaller? ​

Two questions that often get conflated:

  1. How is distillation different from "just training a small model"? Training small from scratch means learning all knowledge from zero; distillation lets the small model inherit knowledge already refined by the large model (soft labels + structural alignment), so at equal capacity it converges better and lands closer to the large model.
  2. Does distillation still matter in the LLM era? Yes, but the form has changed:
    • Data distillation: use a large model offline to generate high-quality data/responses and train a small model (the Alpaca / Self-Instruct approach) — the mainstream route to small LLMs today;
    • Online/offline generation distillation: teacher models guiding students to align on generation tasks (requiring new solutions for problems like distribution collapse);
    • Distilling into stronger small models: MiniLM's teacher-assistant idea continues to pay off in multi-stage distillation.
  3. Distillation ≠ a replacement for quantization: they answer different questions — distillation cuts capacity, quantization cuts bit width. Capacity and bit width are independent dimensions.

8. Shared Limitations ​

  • Training cost comes first: distillation is fundamentally "training cost for inference cost" — without the budget, the payoff shrinks.
  • Capacity gaps: when the teacher-student gap is too large, the student can't keep up; teacher assistants or staged distillation are needed.
  • Task dependence: task-level distillation may generalize poorly to new tasks; generic distillation (the DistilBERT/MiniLM route) is more robust but with milder gains.
  • Beware "teaching to the test": on small test sets, a student can overfit the teacher's noise; evaluate with an independent test set and calibrated distributions.

Further Reading ​

References ​