Skip to content

Knowledge Distillation

At a glance Knowledge distillation (Hinton 2015) lets a large model (teacher) teach a small model (student), transferring "dark knowledge" through soft labels. This article covers the three types of supervision (soft/hard/feature), the role of temperature T, LLM distillation (distilling small models from GPT-4, MiniLLM white-box distillation), and the role of distillation in the "compression trio."

Knowledge Distillation ​

Concept Definition: Passing "Dark Knowledge" to a Smaller Model ​

Knowledge Distillation (KD), systematized by Hinton et al. (2015): a large, high-performing teacher model guides the training of a small student model, so the student learns not only the ground truth but also the teacher's output distribution — which encodes the "dark knowledge" the teacher accumulated over years of training.

Two key insights for understanding KD:

  1. Soft labels carry far more information than hard labels — a hard label says "this is a cat"; a soft label says "cat 0.7, dog 0.2, car 0.001" — the latter tells the student "cats look a bit like dogs and nothing like cars," and this inter-class structure is something hard labels can't teach;
  2. KD is not mere mimicry — it's a denser supervision signal — the teacher's logits, intermediate features, and behavioral trajectories are all "teacher's notes" the student can imitate.

KD, Model Quantization Fundamentals, and Pruning and Sparsification form the model compression trio: quantization cuts bytes, pruning cuts parameters, and distillation shrinks the model itself — the three can be stacked (e.g. distill a smaller model first, then quantize it for deployment).

1. Why Soft Labels Work ​

The Difference in Information Density ​

text
Hard label (one-hot):     [0, 1, 0, 0, 0]                    # 1 bit of information
Soft label (teacher):     [0.01, 0.7, 0.2, 0.08, 0.01]       # inter-class relations, confidence

A hard label only tells the student "the correct answer is class 2"; a soft label also tells it:

  • Class 2 has confidence 0.7 (not absolute — there's ambiguity);
  • Class 3 (0.2) resembles class 2 and is easily confused with it;
  • Class 4 (0.08) is somewhat similar too;
  • Classes 1 and 5 are nothing alike.

These "relative probabilities of the incorrect classes" are the dark knowledge — each sample effectively adds 4 extra "soft supervisors," several times the information of a hard label.

Temperature T: Softening the Distribution ​

A teacher's raw softmax output is often too sharp (one class at 0.99, everything else at 0.001) — such a soft label nearly degenerates into a hard label and carries little information. The temperature T softens the distribution:

text
Plain softmax:      p_i = exp(z_i) / Σ_j exp(z_j)
Softmax with T:     p_i = exp(z_i / T) / Σ_j exp(z_j / T)

The larger T, the flatter the distribution (as T→∞ it degenerates to uniform); T=1 is the plain softmax. Hinton recommends T=2-10 (soften the distribution enough to reveal inter-class relations without flattening it completely).

The Geometric Meaning of Temperature

T is a "thermometer" for the logits — high T means intense thermal motion and a more averaged ("hotter") distribution; low T means a sharper ("colder") distribution. During distillation, the teacher outputs a soft distribution at a high T, and the student learns to mimic it at the same high T; at inference time, the student outputs sharp predictions at T=1.

2. The Three Types of Supervision ​

KD comes in more forms than "mimic the output." By what the student imitates, there are three types:

Supervision TypeWhat the Student LearnsFormApplies To
Soft label (output)The teacher's logit distributionL_student = α·CE(hard) + (1-α)·KL(soft)General
Hard label + soft label mixCorrect answers + inter-class relationsSame as above, α weights the trade-offGeneral
Intermediate featureTeacher's intermediate-layer activationsL = ||feat_s - feat_t||²CV, representation learning
Behavior (trajectory)The teacher's output sequences/decisionsReward shaping in RLReinforcement learning, LLM dialogue

1. Soft-Label KD (Classic Hinton) ​

python
# Simplified KD loss
def kd_loss(student_logits, teacher_logits, labels, T=4, alpha=0.5):
    soft_targets = F.softmax(teacher_logits / T, dim=-1)
    soft_pred = F.log_softmax(student_logits / T, dim=-1)
    kd = F.kl_div(soft_pred, soft_targets, reduction='batchmean') * (T * T)
    ce = F.cross_entropy(student_logits, labels)
    return alpha * kd + (1 - alpha) * ce
  • T*T compensates for the temperature scaling — the KL divergence scales the logits by 1/T, shrinking the loss magnitude by T², so multiplying back by T² brings it to the same scale as CE;
  • α is typically 0.5 (half soft, half hard); raise α for easy tasks and lower it for hard ones.

2. Intermediate-Feature KD (FitNet) ​

Have the student mimic the teacher's intermediate-layer activations — transfer the "representation" too:

text
L_total = L_KD(soft) + β · ||proj(feat_s) - feat_t||²
                                       ↑
                        a projection layer may be needed (the student's dimensions are smaller and must be aligned to the teacher's)
  • Strength: the representation space is compressed — the student learns not just the decision boundary but the feature structure;
  • Best for: student and teacher have different architectures but both do representation tasks (e.g. retrieval).

3. Behavior / Sequence KD ​

Go beyond mimicking output logits — mimic the generation trajectory — have the student learn not just "what the teacher says" but "in what order it says it":

  • Sequence-level KD: the teacher generates the best outputs with beam search, and the student is directly fine-tuned to imitate them (watch out for "exposure bias");
  • Token-level KD: the teacher's next-token distribution at every step serves as the student's soft target;
  • Action-level KD: in reinforcement learning, the teacher's Q-values serve as auxiliary rewards for the student.

3. Knowledge Distillation in the LLM Era ​

LLMs are huge, expensive to train, and closed models (GPT-4) are only accessible via API — which has ironically given distillation a new lease on life:

1. Black-Box Distillation: Using API Models as Teachers ​

Closed LLMs like GPT-4 / Claude don't expose logits — only generated text — so use their outputs as fine-tuning data:

text
Raw prompt → GPT-4 → GPT-4 output
                ↓
                student (7B/13B) fine-tunes on these samples

Typical works:

  • Alpaca: fine-tuned LLaMA 7B on instruction data generated by GPT-3.5 → briefly the strongest open model;
  • Vicuna: fine-tuned LLaMA on GPT-4 conversation data → approaching ChatGPT;
  • WizardLM: distillation of complex instruction data;
  • Tabby: distilling Codex for code scenarios.

Limitation: the teacher's logit distribution is unavailable (black box) — the student can only learn the final outputs, with no dark knowledge — noticeably weaker than white-box distillation.

2. White-Box Distillation: Mimicking with Logits ​

If the teacher is open-source (LLaMA, Qwen, DeepSeek), logits are available — do classic Hinton-style soft-label KD directly:

  • MiniLLM (Gu et al., 2024): reverse KL divergence backpropagation designed for LLMs (plain KL vanishes where the student assigns low probability), using diverse beam search + KL minimization;
  • GKD (Generalized KD): the student learns on trajectories it generates itself, avoiding distribution mismatch;
  • MiniMA: uses "asymmetric KL" (student imitates teacher) to compress LLaMA-2-13B to ~7B scale.

3. Task Distillation: Transferring a Specific Skill ​

The teacher doesn't teach the student everything — it transfers one specific skill:

  • Math distillation: the teacher (e.g. GPT-4 + CoT) generates math reasoning data, and the student learns math (e.g. MetaMath, Mammoth);
  • Tool-use distillation: the teacher demonstrates tool calls, and the student imitates;
  • Code distillation: the teacher generates code + test cases, and the student learns programming;
  • Multimodal distillation: the teacher (e.g. GPT-4V) teaches the student (e.g. LLaVA) visual understanding.

This "narrow domain + strong teacher + lots of synthetic data" pattern has produced a batch of SOTA small models.

Practical Paths for Distilling LLMs

  1. You have the teacher's weights → white-box KD (MiniLLM recipe) → best accuracy;
  2. Only a teacher API → black-box KD (generate data and fine-tune) → simple and viable;
  3. Skill-specific distillation → teacher generates synthetic data → student fine-tunes on that skill;
  4. After distillation, apply Model Quantization Fundamentals + Pruning and Sparsification for deployment.

4. The Compression Trio: Where Distillation Fits ​

Distillation, quantization, and pruning are not competitors — they are complementary:

TechniqueWhat It ChangesTraining CostGainTypical Stacking
DistillationModel architecture (large → small)High (needs a teacher)Smaller model, customizable skillsStep 1: distill 70B down to 7B
QuantizationData format (FP16 → INT4)Low (PTQ takes hours)4× memory, 3-4× speedStep 2: quantize 7B to 4bit
PruningParameter count (remove redundancy)Medium (needs fine-tuning)Even smaller model, sparse speedupStep 3: stack 2:4 sparsity

Suggested engineering order:

text
Pick a teacher → distill a small model → quantize → prune (optional) → deploy

Distill for a smaller architecture first (determines model size) → then quantize to cut bytes (determines deployment speed) → prune on top if needed (extreme compression).

Don't Quantize Before Distilling

Getting the order wrong doubles the work:

  • Quantize first: the teacher's accuracy is compromised, and the distilled student is a defective product;
  • Prune first: pruning destroys the structure, and the student can't learn a stable representation during distillation; Distill first to swap the architecture, then stack quantization/pruning — that's the correct engineering order.

5. Engineering Trade-offs of Distillation ​

QuestionOptionsRules of Thumb
How large a teacherGPT-4 / LLaMA-70B / Qwen-72BBigger is better, but balance against API cost / training resources
How large a student1.5B / 7B / 13BThe deployment target decides: 1.5B for edge, 7B for serving
Data volume5K - 1M examplesData quality > quantity; 100K high-quality samples > 1M noisy samples
Training epochs2-5 epochsEarly stopping to avoid overfitting
α weight0.3-0.7Lower α for easy hard-label tasks; higher α for complex soft-label tasks
Temperature T2-10High T for large models (sharp logits); low T for small models

6. Limitations and Pitfalls of Distillation ​

  1. The capability ceiling is the teacher — the student can't learn what the teacher can't do (if GPT-4 can't do a math problem, the distilled student can't either);
  2. Distribution drift — the student is fine-tuned on teacher outputs, but the inference-time input distribution differs → "exposure bias"; GKD/on-policy distillation helps;
  3. Dark knowledge is undervalued — many distillations use only hard labels (teacher's output text), discarding the logits information — much worse results;
  4. Teacher bias is inherited — teacher biases (political leanings, safety filters) carry over to the student and can't be scrubbed off;
  5. Legal and licensing — many closed APIs' ToS prohibit training models on their outputs (OpenAI ToS explicitly forbids it); check the license before black-box distillation.

The Legal Risks of GPT-4 Distillation

  • OpenAI's ToS explicitly prohibits training competing models on ChatGPT/GPT-4 outputs;
  • Some open datasets (e.g. ShareGPT) were collected from user conversations — a licensing gray zone;
  • Check the ToS before distilling commercial LLM services — or you'll receive a cease-and-desist when you go live.

7. Trade-offs ​

  • Quantization vs. distillation: quantization is cheap, distillation expensive — quantize whenever possible; distill only when you need a smaller model or specific skills;
  • White-box vs. black-box: white-box is more accurate but needs teacher weights; black-box is simple but weaker;
  • Whole-model vs. task distillation: whole-model distillation for general models; task distillation for specific skills (math, code, tools);
  • Distillation vs. training a small model from scratch: distillation usually beats training a same-size student from zero (thanks to teacher supervision);
  • Distillation + LoRA: LoRA + distillation on a quantized model → small model with fine-tuning ability.

Further Reading ​

References ​