Skip to content

Distillation, Pruning, and Low-Rank Factorization

At a glance Beyond quantization, distillation, pruning, and low-rank factorization are three more routes to a smaller model. This article compares how each technique works, where each fits, and how to combine them, so you can pick the right move.

Distillation, Pruning, and Low-Rank Factorization ​

The one-sentence definition: distillation, pruning, and low-rank factorization are three model-slimming routes alongside quantization—distillation has "a small model learn to mimic a large model's outputs," pruning removes "unimportant weights," and low-rank factorization compresses "a weight matrix into two smaller matrices."

Industry insight: quantization is always the first choice in deployment optimization, but it has a ceiling (accuracy loss accelerates below 4-bit, and activation quantization is hard). When quantization isn't enough, or the model is simply too big, the mainstream play is to combine these three with quantization: distill a smaller, more stable student first, then prune or quantize it. Multiple production models at Alibaba, Tencent, and Meta have taken the "distill → quantize" route—distillation handles accuracy, quantization handles speed, each doing its own job. The key is understanding each move's cost-benefit curve and not reaching for a cannon to kill a mosquito.

1. The Model Compression Landscape ​

text
The four model compression techniques
 ├─ Quantization    : lower bit width, saves bandwidth/VRAM, structure untouched   → smallest deployment change
 ├─ Distillation    : small model learns from large model, rebuilds the model itself → needs training, highest ceiling
 ├─ Pruning         : remove unimportant weights/channels, add sparsity            → payoff depends on sparse acceleration support
 └─ Low-rank        : approximate factorization, SVD compression                   → fits fully-connected layers; little gain for convolutions

One-line positioning: quantization compresses precision, distillation compresses capacity, pruning compresses connections, low-rank compresses structure. This article focuses on the last three; for quantization, see Quantization.

2. Knowledge Distillation ​

2.1 The Core Idea: Soft Labels Carry More Information Than Hard Labels ​

Hinton proposed it in 2015's "Distilling the Knowledge in a Neural Network": when training a small model (the student), besides the ground-truth labels, have it learn the soft labels (probability distributions) produced by the large model (the teacher).

text
Large teacher model (already trained)
   │ outputs a softmax probability distribution (softened by temperature T)
   │       ↓ soft labels (e.g., cat 0.6, dog 0.3, bird 0.1)
Small student model (to be trained)
   │ loss = α × cross-entropy(hard labels) + β × KL divergence(soft labels, student outputs)

Why soft labels are powerful: a hard label only tells the model "this is a cat"; a soft label also conveys the inter-class similarity structure—"cats look like dogs but not like birds"—prior knowledge the teacher learned from massive data. So the student learns more with fewer parameters.

2.2 Three Forms of Distillation ​

FormHowTypical scenario
Offline distillation (standard KD)Generate soft labels with the large model first, then train the studentThe classic route, one-time cost
Online distillation (online KD)Teacher and student train together, the student participates in the training dynamicsWhen no ready-made large model exists
Self-distillationA model's deeper layers teach its shallower layersReducing teacher dependence

2.3 Expensive to Train, Best Results ​

Distillation's obvious cost is training: the large model must run inference to generate soft labels (expensive but one-time) plus retraining the student. But the payoff is also the most direct: at the same accuracy target, a distilled student usually beats a same-size model trained from scratch and is naturally more robust. That's why "distill → quantize" is the industry's favorite one-two punch: distillation builds up an accuracy budget, and quantization only eats a slice of it.

3. Pruning ​

3.1 Unstructured Pruning: Sparse but Hard to Accelerate ​

Zero out "unimportant" weights one by one (e.g., by absolute-value threshold), producing a sparse matrix. Compression can be dramatic (50–90% sparse), but ordinary hardware won't speed up—dense matrix math doesn't get faster just because some entries are zero, unless the backend has sparse kernels.

3.2 Structured Pruning: Remove Whole Blocks, Easy to Accelerate ​

Remove channels or entire layers, directly changing tensor shapes—any hardware speeds up (the matrices are simply smaller). The downside: accuracy loss is usually larger than with unstructured pruning.

3.3 Semi-Structured 2:4 Sparsity: Ampere's Sweet Spot ​

Starting with the NVIDIA Ampere architecture, hardware supports 2:4 structured sparsity: keep exactly 2 of every 4 consecutive weights. With dedicated hardware support, INT8/FP16 sparse compute doubles (e.g., A100's 312 FP16 TFLOPS → 624 sparse), and accuracy loss is more controllable than with fully random pruning. The catch: the model must go through 2:4-mask "sparsity-aware training" (similar to QAT) during training/fine-tuning, or the accuracy hit is severe.

Pruning typeCompression gainAccuracy riskHardware speedupEngineering cost
UnstructuredHigh (90% sparse)MediumPoor (almost none)Low (just a mask)
Structured (channels/layers)MediumMedium-highGood (shapes shrink)Medium (needs retraining to recover)
2:4 semi-structuredMedium (fixed 50%)Medium (with training)Good (native tensor core)Medium-high

4. Low-Rank Factorization ​

4.1 The Principle: One Matrix ≈ Two Smaller Matrices ​

Use SVD to factor the weight matrix W (m×n) into U (m×k) × V (k×n); when k << min(m,n), the total parameter count drops from m×n to k×(m+n).

text
Original:   W = U × V     parameters m×n
Compressed: k = 64 (from 4096) → m×k + k×n ≈ 4096×64×2 ≈ 0.5M vs 16.7M, ~97% saved

4.2 Where It Fits: Great for Fully-Connected Layers, Marginal for Convolutions ​

  • Fully-connected / embedding layers: big, low-rank matrices—compression works well;
  • Convolutional layers: already 4D tensors, SVD gains are limited, and the cross-layer structure is complex;
  • LLM FFN layers: mature low-rank methods like LoRA exist, but LoRA was designed for fine-tuning, not for directly compressing an inference model at deployment.

The reality: low-rank factorization is less and less used as a standalone deployment technique—for modern deep models, distillation and quantization deliver bigger gains with less risk. It mostly lives on the training side (LoRA fine-tuning, DeepSeek-V2's MLA, etc.).

5. Side-by-Side Comparison and Combination Strategies ​

5.1 The Comparison Table ​

DimensionDistillationPruningLow-rank factorization
Compression ratioSet by student size, can reach 10–100×30–90% sparse (unstructured)10–100× on fully-connected layers
Accuracy lossLowest (recovered by retraining)Medium (structured pruning recovers with retraining)Medium-high (fixed structural error)
Training costHigh (retrain the student)Medium (may need fine-tuning)Low
Deployment benefitThe model itself gets smaller and fasterStructure gets faster / supports sparsityMatrices get smaller
Hardware dependenceNone2:4 needs Ampere+; unstructured almost noneNone
Ecosystem maturityHigh (plenty of distilled models on Hugging Face)MediumLow (on the deployment side)

5.2 Combination Strategies: Distill First, Then Quantize ​

text
Route A (most common): distill → quantize
   Large model → distill a small model (accuracy up, capacity down) → PTQ/QAT quantization (speed way up)
   Why: distillation makes the model "easier to quantize"; the accuracy quantization eats is covered by distillation's headroom

Route B: prune → quantize
   Large model → 2:4 sparsity or channel pruning → INT8 quantization
   Why: sparsity and quantization stack; on an A100 that's ~4× the throughput of dense FP16

Route C (not recommended): quantize the original model directly
   Easy, but you shoulder the entire accuracy risk

The Real Gains from Combining

On an A100, 2:4 sparsity (2×) + INT8 (2–4×) can theoretically lift throughput 4–8×—ceiling-level "model-layer" optimization in the four-layer framework of Performance Optimization and Capacity Planning.

5.3 Which Technique Fits Which Model ​

  • LLMs: quantization + distillation first (e.g., distill a 1/3-parameter-size LLM); pruning on LLMs is heavily researched but rarely shipped;
  • CV models (ResNet/detection): 2:4 sparsity + quantization is a proven combination;
  • Small/mid NLP models (BERT-class): distilling into TinyBERT/MiniLM is standard practice, then add INT8;
  • Recommender embedding towers: low-rank factorization + quantization is a common pairing.

Trade-offs ​

Decision pointOptionsHow to choose
Compression techniqueQuantization vs distillation vs pruning vs low-rankQuantize by default; add distillation if accuracy falls short; add 2:4 sparsity if hardware supports it
Training resources availableYes → distillation/structured pruning; No → PTQ quantizationNo GPU, no retraining routes
Hardware capabilityAmpere+ → 2:4 sparsity works; older cards → skip sparsityCheck the target card's specs first
Time budget1 week to launch → quantization only; 1 month+ → distillation is viableDistillation retraining takes days

One-line summary: quantization solves "fast," distillation solves "small and accurate," and pruning and low-rank each have their own boundaries—quantize first, distill if that's not enough, add 2:4 sparsity after that; spend your budget in that order.

Further Reading ​

References ​