Appearance
Pruning and Sparsification
Concept Definition: Removing the Parts of the Model That "Don't Pull Their Weight"
Pruning zeroes out or removes unimportant weights or structures in a neural network to make it smaller and faster. Its inspiration comes from the biological brain — synaptic connections are "pruned" during development to improve efficiency, and machine learning can do the same.
Two key insights for understanding pruning:
- Models are often over-parameterized — many weights contribute almost nothing (near 0) to the final output, and removing them barely affects accuracy;
- Sparsity ≠ speedup — zeroing out 50% of weights may not hurt accuracy, but GPUs don't automatically accelerate sparse matrix multiplication. Only structured sparsity + hardware support yields real speedups.
Pruning, Model Quantization Fundamentals, and Knowledge Distillation form the model compression trio. All three have been re-examined in the LLM era — LLMs are huge but expensive to fine-tune, PTQ quantization became the mainstream, and pruning has retreated to a secondary role, though it still earns its place in specific scenarios (edge deployment, extreme compression).
1. Unstructured vs. Structured Pruning
The fundamental divide of pruning — the granularity of zeroing out:
| Type | Zeroing Granularity | Sparsity Rate | Speedup | Best For |
|---|---|---|---|---|
| Unstructured pruning | Individual weights | Still viable at 90%+ | Only with software support | Accuracy-first, research |
| Structured pruning | Whole rows/channels/heads/layers | 30-70% | Native hardware acceleration | Speed-first, deployment |
Unstructured Pruning
Zero out individual weights — the model becomes mathematically sparse (the W matrix full of zeros), but its shape doesn't change — it's still a tensor of the same shape.
text
Original weights: After pruning (unstructured):
1.2 -0.5 0.8 1.2 0 0.8
0.3 0.7 -0.1 0 0.7 0
-0.9 0.2 0.4 -0.9 0 0.4- Strength: excellent accuracy retention (still works after pruning 90%);
- Weakness: GPUs can't accelerate it — sparse matmul needs dedicated kernels that plain cuBLAS doesn't provide; sparse CSR/CSC storage on CUDA even adds metadata overhead and can slow things down.
Structured Pruning
Zero out or remove entire rows/columns/channels/heads at once — the model's shape actually shrinks.
text
Original weights (3×3): After structured pruning (remove column 2) (3×2):
1.2 -0.5 0.8 1.2 0.8
0.3 0.7 -0.1 0.3 -0.1
-0.9 0.2 0.4 -0.9 0.4- Strength: the model truly gets smaller and faster — plain cuBLAS/standard kernels already accelerate it;
- Weakness: larger accuracy loss (remove the wrong channels and accuracy collapses);
- Mainstream directions: channel pruning (CV), head pruning (Transformer), layer pruning (deep networks).
90% Sparse but Not Faster — Why
Many papers report "prune 90% of weights with only 1% accuracy loss," which sounds great — but measured speed is often slower. Reasons:
- Ordinary GPU kernels don't support sparse matmul; dense computation is faster anyway;
- Sparse storage (CSR) carries indirection overhead;
- Quantized sparsity needs to be hardware-aware (e.g. NVIDIA 2:4 sparse Tensor Cores — see below) to actually speed up. The "speedups" in pruning papers are often software simulations — be skeptical in production engineering.
2. The Lottery Ticket Hypothesis
The Lottery Ticket Hypothesis (Frankle & Carbin, 2018): inside a trained dense network hides a sparse "winning ticket" — as long as this subnetwork's initial weights match the original network's, training it alone can reach the original network's accuracy.
text
Train the network → prune (keep the mask) → reset unkept weights to initial values → retrain
↑
Still reaches the original accuracy → a winning ticketImplications:
- It proves models are over-parameterized — a large fraction of weights are "redundant backups";
- It reframes pruning — not "chop off the bad ones after training," but "find the subnetwork that already won at initialization";
- Extension: late resetting (keep weights after some training steps instead of the exact initial weights) lets large models find winning tickets too.
But the Lottery Ticket Hypothesis has cooled in the LLM era — LLMs are too large, and the "train → prune → retrain" cycle is too expensive; searching for subnetworks is nearly infeasible.
3. The Pruning Algorithm Family
1. Magnitude Pruning
The simplest — sort weights by absolute value and zero out the small ones. Rationale: small weights contribute little to the output.
text
score(w) = |w| # weight magnitude
threshold = percentile(|w|, sparsity_ratio)
mask = (|w| > threshold)
W_pruned = W * mask- Strength: simple to implement, no gradient information needed;
- Weakness: large magnitude ≠ importance (a small weight with large gradient can also be critical);
- Status: the baseline and still the industry's default starting point.
2. Iterative Pruning
Pruning too much at once collapses accuracy — prune in multiple rounds instead:
text
Train → prune 10% → fine-tune to recover → prune 10% → fine-tune to recover → ... → reach the target sparsityEach round of fine-tuning lets the model redistribute importance among the "remaining weights" — far better accuracy than pruning 50% in one shot.
3. Taylor Pruning (First-Order Importance)
Use gradient information to assess weight importance:
text
score(w) = |w · ∂L/∂w| # weight × gradient, approximating "how much the loss rises if this is removed"Rationale: the loss change from removing weight w ≈ w · ∂L/∂w (first-order Taylor expansion). More accurate than magnitude, but requires forward and backward passes for gradients.
4. Movement Pruning
Not based on the "current magnitude" but on the "trend of magnitude during training":
text
score(w) = Σ_t w_t · ∂L/∂w_t # accumulated: "drifting toward zero or away from it"Drifting further from 0 → important (keep); drifting closer to 0 → unimportant (prune). Suited for pruning during fine-tuning (e.g. pruning while fine-tuning BERT).
Quick Algorithm Selection Guide
- CV models: magnitude pruning + iterative fine-tuning — mature in practice;
- Transformer fine-tuning: Movement Pruning (fits fine-tuning);
- LLMs: magnitude + activation-aware (see Wanda);
- Chasing extreme sparsity: Taylor + iterative + Lottery Ticket.
4. NVIDIA 2:4 Structured Sparsity
Hardware Support
Tensor Cores since NVIDIA Ampere (A100) support 2:4 structured sparsity (also called 1:2 sparsity): in every group of 4 consecutive weights, exactly 2 are zero — and the hardware accelerates matmul by 2×.
text
Dense weights (in groups of 4):
[1.2 0.5 0.8 0.3] [0.4 0.9 0.1 0.7] ...
After 2:4 sparsification (2 nonzeros per group):
[1.2 0 0.8 0 ] [0 0.9 0 0.7] ...
↑
Tensor Cores accelerate with dedicated instructions: matmul 2×- Strength: genuine 2× hardware acceleration (A100/H100 Sparse Tensor Cores);
- Limitation: the sparsity rate is fixed at 50% (2/4 = 50%) — not tunable;
- Tooling: NVIDIA
sparsetrainer, PyTorch 2:4 sparsity support.
The Best Combo for 2:4 Sparsity
2:4 sparsity + INT4 quantization: first use magnitude pruning to reach 2:4 sparsity (~1% loss), then INT4 quantization (~1% loss) — overall that's 4× memory (quantization) + 2× compute (sparsity) = 8× total gain at ~2% accuracy loss. On H100, this is an extremely cost-effective path to extreme compression.
Other Structured Sparsity Formats
- 1:2 / 1:4 / 1:8: higher sparsity but little hardware support;
- Block sparsity: sparsity in blocks (used by the sparse variant of FlashAttention);
- Head sparsity: prune entire attention heads — viable on LLMs but accuracy risk is high.
5. Pruning in the LLM Era
Traditional pruning runs into trouble on LLMs:
- Models are large — the "prune → fine-tune" loop is too expensive;
- Fine-tuning data is closed (the GPT series can't be retrained);
- Structured pruning destroys pretrained representations, and accuracy degrades badly.
Hence LLM-specific pruning methods:
Wanda (Pruning by Weights and Activations)
- Idea: assess weight importance not just by magnitude — combine it with activation magnitude: weights with large
|w| · ||x||matter; - Strength: no retraining or fine-tuning needed — pure forward-pass pruning; a 70B model takes hours;
- Sparsity: 50% unstructured with <2% accuracy loss;
- Limitation: unstructured → real speedups require 2:4 sparse rearrangement.
LLM-Pruner
- Idea: relies on structured pruning + LoRA fine-tuning for recovery;
- Flow: find dependency groups → structured pruning → LoRA fine-tune for a few hundred steps → accuracy recovered;
- Sparsity: 20-30% structured with <5% loss;
- Strength: prunes whole heads/MLP dimensions — hardware-acceleratable.
ShortGPT (Layer Pruning)
- Idea: remove redundant entire layers from the Transformer — some layers have outputs nearly identical to their inputs and can be skipped;
- Sparsity: remove 25% of layers with <2% accuracy loss (validated on LLaMA-2-70B);
- Strength: removing whole layers → shallower model → real speedup;
- Limitation: remove the wrong layers and accuracy collapses — inter-layer similarity must be computed to pick which ones.
The Practical Value of LLM Pruning Is Limited
In the LLM era, PTQ quantization has largely replaced pruning:
- Quantization (W4A16) compresses 4× with ~1-3% accuracy loss and no training;
- After 50% pruning, LoRA fine-tuning is usually needed to recover accuracy — and the compression ratio may still lose to quantization. When LLM pruning still makes sense:
- Extremely tight memory (edge deployment): quantization + sparsity stacked;
- Extreme compute (H100 sparse Tensor Cores): 2:4 sparsity + INT4;
- Whole-layer removal (ShortGPT): a shallower model gives real speedups. For ordinary online LLM serving, default to Weight-Only Quantization and Mixed Precision and treat pruning as a complement.
6. Sparsification and Acceleration Paths
Different sparsity rates call for different acceleration paths:
| Sparsity | Sparsity Type | Acceleration Method | Best For |
|---|---|---|---|
| 50% (2:4) | Structured | NVIDIA Sparse Tensor Cores | H100/A100, 2× speedup |
| 70-90% | Unstructured | Dedicated sparse kernels (e.g. DeepSparse) | CPU edge deployment |
| 90%+ | Extreme | Software simulation | Research |
| Whole-layer removal | Structured | Shallower model | Any hardware |
| Head pruning | Structured | Smaller model | Any hardware |
7. Trade-offs
- Quantization vs. pruning: quantization first for LLM inference (cheaper), pruning as a complement; comparable for CV models;
- Structured vs. unstructured: deployment favors structured (real speedups); research favors unstructured for accuracy;
- Magnitude vs. Taylor: accuracy-first picks Taylor; simple deployment picks magnitude;
- LLM scenarios: PTQ quantization leads; stack 2:4 sparsity on top (quantization + sparsity = 8× compression);
- Training resources: if you can fine-tune, use iterative + LoRA recovery; if not, use Wanda (pure forward pass).
Further Reading
- Model Quantization Fundamentals — the "byte-cutting" tool of the compression trio
- Weight-Only Quantization and Mixed Precision — the engineering path for stacking with pruning
- Knowledge Distillation — the "training-side" weapon of the compression trio
- Kernel Fusion and Custom Kernels — kernel implementation of sparse matmul
- GPU Architecture and Optimization — hardware details of Sparse Tensor Cores
- llama.cpp and GGUF — sparse inference on CPU
- Classic Papers in Depth — the Lottery Ticket/Wanda/ShortGPT papers
References
- Frankle & Carbin. The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks (ICLR 2019) — the original Lottery Ticket paper
- Han et al. Learning both Weights and Connections for Efficient Neural Networks (NeurIPS 2015) — classic magnitude pruning
- Mishra et al. Movement Pruning: Adaptive Sparsity by Fine-Tuning (NeurIPS 2020) — the Movement Pruning paper
- Sun et al. Wanda: Pruning LLMs by Weights and Activations (ICLR 2024) — pure forward-pass pruning for LLMs
- Ma et al. LLM-Pruner: On the Structural Pruning of Large Language Models (NeurIPS 2023) — structured pruning for LLMs
- Men et al. ShortGPT: Layers in Large Language Models are More Redundant Than You Expect (2024) — LLM layer pruning
- NVIDIA 2:4 Structured Sparsity Documentation — official documentation of 2:4 hardware sparsity
- Kwon et al. Sparse GPU Kernels for Masking and Point-Error-Correcting (SC23) — engineering implementation of sparse kernels