Appearance
Distillation, Pruning, and Low-Rank Factorization
The one-sentence definition: distillation, pruning, and low-rank factorization are three model-slimming routes alongside quantization—distillation has "a small model learn to mimic a large model's outputs," pruning removes "unimportant weights," and low-rank factorization compresses "a weight matrix into two smaller matrices."
Industry insight: quantization is always the first choice in deployment optimization, but it has a ceiling (accuracy loss accelerates below 4-bit, and activation quantization is hard). When quantization isn't enough, or the model is simply too big, the mainstream play is to combine these three with quantization: distill a smaller, more stable student first, then prune or quantize it. Multiple production models at Alibaba, Tencent, and Meta have taken the "distill → quantize" route—distillation handles accuracy, quantization handles speed, each doing its own job. The key is understanding each move's cost-benefit curve and not reaching for a cannon to kill a mosquito.
1. The Model Compression Landscape
text
The four model compression techniques
├─ Quantization : lower bit width, saves bandwidth/VRAM, structure untouched → smallest deployment change
├─ Distillation : small model learns from large model, rebuilds the model itself → needs training, highest ceiling
├─ Pruning : remove unimportant weights/channels, add sparsity → payoff depends on sparse acceleration support
└─ Low-rank : approximate factorization, SVD compression → fits fully-connected layers; little gain for convolutionsOne-line positioning: quantization compresses precision, distillation compresses capacity, pruning compresses connections, low-rank compresses structure. This article focuses on the last three; for quantization, see Quantization.
2. Knowledge Distillation
2.1 The Core Idea: Soft Labels Carry More Information Than Hard Labels
Hinton proposed it in 2015's "Distilling the Knowledge in a Neural Network": when training a small model (the student), besides the ground-truth labels, have it learn the soft labels (probability distributions) produced by the large model (the teacher).
text
Large teacher model (already trained)
│ outputs a softmax probability distribution (softened by temperature T)
│ ↓ soft labels (e.g., cat 0.6, dog 0.3, bird 0.1)
Small student model (to be trained)
│ loss = α × cross-entropy(hard labels) + β × KL divergence(soft labels, student outputs)Why soft labels are powerful: a hard label only tells the model "this is a cat"; a soft label also conveys the inter-class similarity structure—"cats look like dogs but not like birds"—prior knowledge the teacher learned from massive data. So the student learns more with fewer parameters.
2.2 Three Forms of Distillation
| Form | How | Typical scenario |
|---|---|---|
| Offline distillation (standard KD) | Generate soft labels with the large model first, then train the student | The classic route, one-time cost |
| Online distillation (online KD) | Teacher and student train together, the student participates in the training dynamics | When no ready-made large model exists |
| Self-distillation | A model's deeper layers teach its shallower layers | Reducing teacher dependence |
2.3 Expensive to Train, Best Results
Distillation's obvious cost is training: the large model must run inference to generate soft labels (expensive but one-time) plus retraining the student. But the payoff is also the most direct: at the same accuracy target, a distilled student usually beats a same-size model trained from scratch and is naturally more robust. That's why "distill → quantize" is the industry's favorite one-two punch: distillation builds up an accuracy budget, and quantization only eats a slice of it.
3. Pruning
3.1 Unstructured Pruning: Sparse but Hard to Accelerate
Zero out "unimportant" weights one by one (e.g., by absolute-value threshold), producing a sparse matrix. Compression can be dramatic (50–90% sparse), but ordinary hardware won't speed up—dense matrix math doesn't get faster just because some entries are zero, unless the backend has sparse kernels.
3.2 Structured Pruning: Remove Whole Blocks, Easy to Accelerate
Remove channels or entire layers, directly changing tensor shapes—any hardware speeds up (the matrices are simply smaller). The downside: accuracy loss is usually larger than with unstructured pruning.
3.3 Semi-Structured 2:4 Sparsity: Ampere's Sweet Spot
Starting with the NVIDIA Ampere architecture, hardware supports 2:4 structured sparsity: keep exactly 2 of every 4 consecutive weights. With dedicated hardware support, INT8/FP16 sparse compute doubles (e.g., A100's 312 FP16 TFLOPS → 624 sparse), and accuracy loss is more controllable than with fully random pruning. The catch: the model must go through 2:4-mask "sparsity-aware training" (similar to QAT) during training/fine-tuning, or the accuracy hit is severe.
| Pruning type | Compression gain | Accuracy risk | Hardware speedup | Engineering cost |
|---|---|---|---|---|
| Unstructured | High (90% sparse) | Medium | Poor (almost none) | Low (just a mask) |
| Structured (channels/layers) | Medium | Medium-high | Good (shapes shrink) | Medium (needs retraining to recover) |
| 2:4 semi-structured | Medium (fixed 50%) | Medium (with training) | Good (native tensor core) | Medium-high |
4. Low-Rank Factorization
4.1 The Principle: One Matrix ≈ Two Smaller Matrices
Use SVD to factor the weight matrix W (m×n) into U (m×k) × V (k×n); when k << min(m,n), the total parameter count drops from m×n to k×(m+n).
text
Original: W = U × V parameters m×n
Compressed: k = 64 (from 4096) → m×k + k×n ≈ 4096×64×2 ≈ 0.5M vs 16.7M, ~97% saved4.2 Where It Fits: Great for Fully-Connected Layers, Marginal for Convolutions
- Fully-connected / embedding layers: big, low-rank matrices—compression works well;
- Convolutional layers: already 4D tensors, SVD gains are limited, and the cross-layer structure is complex;
- LLM FFN layers: mature low-rank methods like LoRA exist, but LoRA was designed for fine-tuning, not for directly compressing an inference model at deployment.
The reality: low-rank factorization is less and less used as a standalone deployment technique—for modern deep models, distillation and quantization deliver bigger gains with less risk. It mostly lives on the training side (LoRA fine-tuning, DeepSeek-V2's MLA, etc.).
5. Side-by-Side Comparison and Combination Strategies
5.1 The Comparison Table
| Dimension | Distillation | Pruning | Low-rank factorization |
|---|---|---|---|
| Compression ratio | Set by student size, can reach 10–100× | 30–90% sparse (unstructured) | 10–100× on fully-connected layers |
| Accuracy loss | Lowest (recovered by retraining) | Medium (structured pruning recovers with retraining) | Medium-high (fixed structural error) |
| Training cost | High (retrain the student) | Medium (may need fine-tuning) | Low |
| Deployment benefit | The model itself gets smaller and faster | Structure gets faster / supports sparsity | Matrices get smaller |
| Hardware dependence | None | 2:4 needs Ampere+; unstructured almost none | None |
| Ecosystem maturity | High (plenty of distilled models on Hugging Face) | Medium | Low (on the deployment side) |
5.2 Combination Strategies: Distill First, Then Quantize
text
Route A (most common): distill → quantize
Large model → distill a small model (accuracy up, capacity down) → PTQ/QAT quantization (speed way up)
Why: distillation makes the model "easier to quantize"; the accuracy quantization eats is covered by distillation's headroom
Route B: prune → quantize
Large model → 2:4 sparsity or channel pruning → INT8 quantization
Why: sparsity and quantization stack; on an A100 that's ~4× the throughput of dense FP16
Route C (not recommended): quantize the original model directly
Easy, but you shoulder the entire accuracy riskThe Real Gains from Combining
On an A100, 2:4 sparsity (2×) + INT8 (2–4×) can theoretically lift throughput 4–8×—ceiling-level "model-layer" optimization in the four-layer framework of Performance Optimization and Capacity Planning.
5.3 Which Technique Fits Which Model
- LLMs: quantization + distillation first (e.g., distill a 1/3-parameter-size LLM); pruning on LLMs is heavily researched but rarely shipped;
- CV models (ResNet/detection): 2:4 sparsity + quantization is a proven combination;
- Small/mid NLP models (BERT-class): distilling into TinyBERT/MiniLM is standard practice, then add INT8;
- Recommender embedding towers: low-rank factorization + quantization is a common pairing.
Trade-offs
| Decision point | Options | How to choose |
|---|---|---|
| Compression technique | Quantization vs distillation vs pruning vs low-rank | Quantize by default; add distillation if accuracy falls short; add 2:4 sparsity if hardware supports it |
| Training resources available | Yes → distillation/structured pruning; No → PTQ quantization | No GPU, no retraining routes |
| Hardware capability | Ampere+ → 2:4 sparsity works; older cards → skip sparsity | Check the target card's specs first |
| Time budget | 1 week to launch → quantization only; 1 month+ → distillation is viable | Distillation retraining takes days |
One-line summary: quantization solves "fast," distillation solves "small and accurate," and pruning and low-rank each have their own boundaries—quantize first, distill if that's not enough, add 2:4 sparsity after that; spend your budget in that order.
Further Reading
- Quantization — the smallest deployment change and most direct payoff of the four techniques
- Papers: Distillation Classics — KD and its family, paper by paper
- Model Optimization in Practice — concrete steps and acceptance for combination strategies
- GPUs and Hardware Selection — the hardware prerequisites for 2:4 sparsity
- LLM Inference Optimization — combining compression and inference for LLMs
References
- Distilling the Knowledge in a Neural Network (Hinton et al., 2015)
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour (includes structured pruning discussion)
- NVIDIA 2:4 Structured Sparsity whitepaper
- A Survey of Model Compression and Acceleration for Deep Neural Networks (Cheng et al.)
- Hugging Face distilled model collection (TinyBERT, DistilBERT, etc.)