Appearance
Knowledge Distillation
Concept Definition: Passing "Dark Knowledge" to a Smaller Model
Knowledge Distillation (KD), systematized by Hinton et al. (2015): a large, high-performing teacher model guides the training of a small student model, so the student learns not only the ground truth but also the teacher's output distribution — which encodes the "dark knowledge" the teacher accumulated over years of training.
Two key insights for understanding KD:
- Soft labels carry far more information than hard labels — a hard label says "this is a cat"; a soft label says "cat 0.7, dog 0.2, car 0.001" — the latter tells the student "cats look a bit like dogs and nothing like cars," and this inter-class structure is something hard labels can't teach;
- KD is not mere mimicry — it's a denser supervision signal — the teacher's logits, intermediate features, and behavioral trajectories are all "teacher's notes" the student can imitate.
KD, Model Quantization Fundamentals, and Pruning and Sparsification form the model compression trio: quantization cuts bytes, pruning cuts parameters, and distillation shrinks the model itself — the three can be stacked (e.g. distill a smaller model first, then quantize it for deployment).
1. Why Soft Labels Work
The Difference in Information Density
text
Hard label (one-hot): [0, 1, 0, 0, 0] # 1 bit of information
Soft label (teacher): [0.01, 0.7, 0.2, 0.08, 0.01] # inter-class relations, confidenceA hard label only tells the student "the correct answer is class 2"; a soft label also tells it:
- Class 2 has confidence 0.7 (not absolute — there's ambiguity);
- Class 3 (0.2) resembles class 2 and is easily confused with it;
- Class 4 (0.08) is somewhat similar too;
- Classes 1 and 5 are nothing alike.
These "relative probabilities of the incorrect classes" are the dark knowledge — each sample effectively adds 4 extra "soft supervisors," several times the information of a hard label.
Temperature T: Softening the Distribution
A teacher's raw softmax output is often too sharp (one class at 0.99, everything else at 0.001) — such a soft label nearly degenerates into a hard label and carries little information. The temperature T softens the distribution:
text
Plain softmax: p_i = exp(z_i) / Σ_j exp(z_j)
Softmax with T: p_i = exp(z_i / T) / Σ_j exp(z_j / T)The larger T, the flatter the distribution (as T→∞ it degenerates to uniform); T=1 is the plain softmax. Hinton recommends T=2-10 (soften the distribution enough to reveal inter-class relations without flattening it completely).
The Geometric Meaning of Temperature
T is a "thermometer" for the logits — high T means intense thermal motion and a more averaged ("hotter") distribution; low T means a sharper ("colder") distribution. During distillation, the teacher outputs a soft distribution at a high T, and the student learns to mimic it at the same high T; at inference time, the student outputs sharp predictions at T=1.
2. The Three Types of Supervision
KD comes in more forms than "mimic the output." By what the student imitates, there are three types:
| Supervision Type | What the Student Learns | Form | Applies To |
|---|---|---|---|
| Soft label (output) | The teacher's logit distribution | L_student = α·CE(hard) + (1-α)·KL(soft) | General |
| Hard label + soft label mix | Correct answers + inter-class relations | Same as above, α weights the trade-off | General |
| Intermediate feature | Teacher's intermediate-layer activations | L = ||feat_s - feat_t||² | CV, representation learning |
| Behavior (trajectory) | The teacher's output sequences/decisions | Reward shaping in RL | Reinforcement learning, LLM dialogue |
1. Soft-Label KD (Classic Hinton)
python
# Simplified KD loss
def kd_loss(student_logits, teacher_logits, labels, T=4, alpha=0.5):
soft_targets = F.softmax(teacher_logits / T, dim=-1)
soft_pred = F.log_softmax(student_logits / T, dim=-1)
kd = F.kl_div(soft_pred, soft_targets, reduction='batchmean') * (T * T)
ce = F.cross_entropy(student_logits, labels)
return alpha * kd + (1 - alpha) * ceT*Tcompensates for the temperature scaling — the KL divergence scales the logits by 1/T, shrinking the loss magnitude by T², so multiplying back by T² brings it to the same scale as CE;- α is typically 0.5 (half soft, half hard); raise α for easy tasks and lower it for hard ones.
2. Intermediate-Feature KD (FitNet)
Have the student mimic the teacher's intermediate-layer activations — transfer the "representation" too:
text
L_total = L_KD(soft) + β · ||proj(feat_s) - feat_t||²
↑
a projection layer may be needed (the student's dimensions are smaller and must be aligned to the teacher's)- Strength: the representation space is compressed — the student learns not just the decision boundary but the feature structure;
- Best for: student and teacher have different architectures but both do representation tasks (e.g. retrieval).
3. Behavior / Sequence KD
Go beyond mimicking output logits — mimic the generation trajectory — have the student learn not just "what the teacher says" but "in what order it says it":
- Sequence-level KD: the teacher generates the best outputs with beam search, and the student is directly fine-tuned to imitate them (watch out for "exposure bias");
- Token-level KD: the teacher's next-token distribution at every step serves as the student's soft target;
- Action-level KD: in reinforcement learning, the teacher's Q-values serve as auxiliary rewards for the student.
3. Knowledge Distillation in the LLM Era
LLMs are huge, expensive to train, and closed models (GPT-4) are only accessible via API — which has ironically given distillation a new lease on life:
1. Black-Box Distillation: Using API Models as Teachers
Closed LLMs like GPT-4 / Claude don't expose logits — only generated text — so use their outputs as fine-tuning data:
text
Raw prompt → GPT-4 → GPT-4 output
↓
student (7B/13B) fine-tunes on these samplesTypical works:
- Alpaca: fine-tuned LLaMA 7B on instruction data generated by GPT-3.5 → briefly the strongest open model;
- Vicuna: fine-tuned LLaMA on GPT-4 conversation data → approaching ChatGPT;
- WizardLM: distillation of complex instruction data;
- Tabby: distilling Codex for code scenarios.
Limitation: the teacher's logit distribution is unavailable (black box) — the student can only learn the final outputs, with no dark knowledge — noticeably weaker than white-box distillation.
2. White-Box Distillation: Mimicking with Logits
If the teacher is open-source (LLaMA, Qwen, DeepSeek), logits are available — do classic Hinton-style soft-label KD directly:
- MiniLLM (Gu et al., 2024): reverse KL divergence backpropagation designed for LLMs (plain KL vanishes where the student assigns low probability), using diverse beam search + KL minimization;
- GKD (Generalized KD): the student learns on trajectories it generates itself, avoiding distribution mismatch;
- MiniMA: uses "asymmetric KL" (student imitates teacher) to compress LLaMA-2-13B to ~7B scale.
3. Task Distillation: Transferring a Specific Skill
The teacher doesn't teach the student everything — it transfers one specific skill:
- Math distillation: the teacher (e.g. GPT-4 + CoT) generates math reasoning data, and the student learns math (e.g. MetaMath, Mammoth);
- Tool-use distillation: the teacher demonstrates tool calls, and the student imitates;
- Code distillation: the teacher generates code + test cases, and the student learns programming;
- Multimodal distillation: the teacher (e.g. GPT-4V) teaches the student (e.g. LLaVA) visual understanding.
This "narrow domain + strong teacher + lots of synthetic data" pattern has produced a batch of SOTA small models.
Practical Paths for Distilling LLMs
- You have the teacher's weights → white-box KD (MiniLLM recipe) → best accuracy;
- Only a teacher API → black-box KD (generate data and fine-tune) → simple and viable;
- Skill-specific distillation → teacher generates synthetic data → student fine-tunes on that skill;
- After distillation, apply Model Quantization Fundamentals + Pruning and Sparsification for deployment.
4. The Compression Trio: Where Distillation Fits
Distillation, quantization, and pruning are not competitors — they are complementary:
| Technique | What It Changes | Training Cost | Gain | Typical Stacking |
|---|---|---|---|---|
| Distillation | Model architecture (large → small) | High (needs a teacher) | Smaller model, customizable skills | Step 1: distill 70B down to 7B |
| Quantization | Data format (FP16 → INT4) | Low (PTQ takes hours) | 4× memory, 3-4× speed | Step 2: quantize 7B to 4bit |
| Pruning | Parameter count (remove redundancy) | Medium (needs fine-tuning) | Even smaller model, sparse speedup | Step 3: stack 2:4 sparsity |
Suggested engineering order:
text
Pick a teacher → distill a small model → quantize → prune (optional) → deployDistill for a smaller architecture first (determines model size) → then quantize to cut bytes (determines deployment speed) → prune on top if needed (extreme compression).
Don't Quantize Before Distilling
Getting the order wrong doubles the work:
- Quantize first: the teacher's accuracy is compromised, and the distilled student is a defective product;
- Prune first: pruning destroys the structure, and the student can't learn a stable representation during distillation; Distill first to swap the architecture, then stack quantization/pruning — that's the correct engineering order.
5. Engineering Trade-offs of Distillation
| Question | Options | Rules of Thumb |
|---|---|---|
| How large a teacher | GPT-4 / LLaMA-70B / Qwen-72B | Bigger is better, but balance against API cost / training resources |
| How large a student | 1.5B / 7B / 13B | The deployment target decides: 1.5B for edge, 7B for serving |
| Data volume | 5K - 1M examples | Data quality > quantity; 100K high-quality samples > 1M noisy samples |
| Training epochs | 2-5 epochs | Early stopping to avoid overfitting |
| α weight | 0.3-0.7 | Lower α for easy hard-label tasks; higher α for complex soft-label tasks |
| Temperature T | 2-10 | High T for large models (sharp logits); low T for small models |
6. Limitations and Pitfalls of Distillation
- The capability ceiling is the teacher — the student can't learn what the teacher can't do (if GPT-4 can't do a math problem, the distilled student can't either);
- Distribution drift — the student is fine-tuned on teacher outputs, but the inference-time input distribution differs → "exposure bias"; GKD/on-policy distillation helps;
- Dark knowledge is undervalued — many distillations use only hard labels (teacher's output text), discarding the logits information — much worse results;
- Teacher bias is inherited — teacher biases (political leanings, safety filters) carry over to the student and can't be scrubbed off;
- Legal and licensing — many closed APIs' ToS prohibit training models on their outputs (OpenAI ToS explicitly forbids it); check the license before black-box distillation.
The Legal Risks of GPT-4 Distillation
- OpenAI's ToS explicitly prohibits training competing models on ChatGPT/GPT-4 outputs;
- Some open datasets (e.g. ShareGPT) were collected from user conversations — a licensing gray zone;
- Check the ToS before distilling commercial LLM services — or you'll receive a cease-and-desist when you go live.
7. Trade-offs
- Quantization vs. distillation: quantization is cheap, distillation expensive — quantize whenever possible; distill only when you need a smaller model or specific skills;
- White-box vs. black-box: white-box is more accurate but needs teacher weights; black-box is simple but weaker;
- Whole-model vs. task distillation: whole-model distillation for general models; task distillation for specific skills (math, code, tools);
- Distillation vs. training a small model from scratch: distillation usually beats training a same-size student from zero (thanks to teacher supervision);
- Distillation + LoRA: LoRA + distillation on a quantized model → small model with fine-tuning ability.
Further Reading
- Pruning and Sparsification — the "parameter-cutting" tool of the compression trio
- Model Quantization Fundamentals — the "byte-cutting" tool of the compression trio
- Weight-Only Quantization and Mixed Precision — deployment acceleration for the distilled small model
- Model Serving and Orchestration — deploying and orchestrating small models
- Tuning and Performance Optimization — the practical path of LoRA + distillation
- Classic Papers in Depth — links to the KD/MiniLLM/Alpaca papers
- Frontier Advances — the latest progress in distillation algorithms
References
- Hinton, Vinyals, Dean. Distilling the Knowledge in a Neural Network (NeurIPS Workshop 2015) — the foundational paper of knowledge distillation
- Romero et al. FitNets: Hints for Thin Deep Nets (ICLR 2015) — intermediate-layer distillation
- Sanh et al. DistilBERT: A distilled version of BERT (2019) — the engineering exemplar of BERT distillation
- Gu et al. MiniLLM: Knowledge Distillation of Large Language Models (ICLR 2024) — white-box distillation for LLMs
- Taori et al. Alpaca: A Strong, Replicable Instruction-Following Model (2023) — the classic of black-box LLM distillation
- Mukherjee et al. SeqKD: Systematic Generalization Distillation — sequence-level KD
- Agarwal et al. On-Policy Distillation of Language Models (2024) — GKD/on-policy KD
- Wang et al. MetaMath: Bootstrap Your Mathematical Questions (2024) — an exemplar of task distillation