Skip to content

Pruning and Sparsification

At a glance Pruning removes unimportant weights or structures to make models smaller and faster. This article covers unstructured vs. structured pruning, the Lottery Ticket Hypothesis, magnitude/iterative/Taylor/movement pruning algorithms, NVIDIA 2:4 structured sparsity, and Wanda/LLM-Pruner in the LLM era.

Pruning and Sparsification ​

Concept Definition: Removing the Parts of the Model That "Don't Pull Their Weight" ​

Pruning zeroes out or removes unimportant weights or structures in a neural network to make it smaller and faster. Its inspiration comes from the biological brain — synaptic connections are "pruned" during development to improve efficiency, and machine learning can do the same.

Two key insights for understanding pruning:

  1. Models are often over-parameterized — many weights contribute almost nothing (near 0) to the final output, and removing them barely affects accuracy;
  2. Sparsity ≠ speedup — zeroing out 50% of weights may not hurt accuracy, but GPUs don't automatically accelerate sparse matrix multiplication. Only structured sparsity + hardware support yields real speedups.

Pruning, Model Quantization Fundamentals, and Knowledge Distillation form the model compression trio. All three have been re-examined in the LLM era — LLMs are huge but expensive to fine-tune, PTQ quantization became the mainstream, and pruning has retreated to a secondary role, though it still earns its place in specific scenarios (edge deployment, extreme compression).

1. Unstructured vs. Structured Pruning ​

The fundamental divide of pruning — the granularity of zeroing out:

TypeZeroing GranularitySparsity RateSpeedupBest For
Unstructured pruningIndividual weightsStill viable at 90%+Only with software supportAccuracy-first, research
Structured pruningWhole rows/channels/heads/layers30-70%Native hardware accelerationSpeed-first, deployment

Unstructured Pruning ​

Zero out individual weights — the model becomes mathematically sparse (the W matrix full of zeros), but its shape doesn't change — it's still a tensor of the same shape.

text
Original weights:     After pruning (unstructured):
1.2  -0.5  0.8        1.2   0    0.8
0.3   0.7 -0.1         0   0.7   0
-0.9  0.2  0.4        -0.9   0   0.4
  • Strength: excellent accuracy retention (still works after pruning 90%);
  • Weakness: GPUs can't accelerate it — sparse matmul needs dedicated kernels that plain cuBLAS doesn't provide; sparse CSR/CSC storage on CUDA even adds metadata overhead and can slow things down.

Structured Pruning ​

Zero out or remove entire rows/columns/channels/heads at once — the model's shape actually shrinks.

text
Original weights (3×3):   After structured pruning (remove column 2) (3×2):
1.2  -0.5  0.8            1.2  0.8
0.3   0.7 -0.1            0.3 -0.1
-0.9  0.2  0.4            -0.9  0.4
  • Strength: the model truly gets smaller and faster — plain cuBLAS/standard kernels already accelerate it;
  • Weakness: larger accuracy loss (remove the wrong channels and accuracy collapses);
  • Mainstream directions: channel pruning (CV), head pruning (Transformer), layer pruning (deep networks).

90% Sparse but Not Faster — Why

Many papers report "prune 90% of weights with only 1% accuracy loss," which sounds great — but measured speed is often slower. Reasons:

  1. Ordinary GPU kernels don't support sparse matmul; dense computation is faster anyway;
  2. Sparse storage (CSR) carries indirection overhead;
  3. Quantized sparsity needs to be hardware-aware (e.g. NVIDIA 2:4 sparse Tensor Cores — see below) to actually speed up. The "speedups" in pruning papers are often software simulations — be skeptical in production engineering.

2. The Lottery Ticket Hypothesis ​

The Lottery Ticket Hypothesis (Frankle & Carbin, 2018): inside a trained dense network hides a sparse "winning ticket" — as long as this subnetwork's initial weights match the original network's, training it alone can reach the original network's accuracy.

text
Train the network → prune (keep the mask) → reset unkept weights to initial values → retrain
                                                                        ↑
                                                        Still reaches the original accuracy → a winning ticket

Implications:

  1. It proves models are over-parameterized — a large fraction of weights are "redundant backups";
  2. It reframes pruning — not "chop off the bad ones after training," but "find the subnetwork that already won at initialization";
  3. Extension: late resetting (keep weights after some training steps instead of the exact initial weights) lets large models find winning tickets too.

But the Lottery Ticket Hypothesis has cooled in the LLM era — LLMs are too large, and the "train → prune → retrain" cycle is too expensive; searching for subnetworks is nearly infeasible.

3. The Pruning Algorithm Family ​

1. Magnitude Pruning ​

The simplest — sort weights by absolute value and zero out the small ones. Rationale: small weights contribute little to the output.

text
score(w) = |w|   # weight magnitude
threshold = percentile(|w|, sparsity_ratio)
mask = (|w| > threshold)
W_pruned = W * mask
  • Strength: simple to implement, no gradient information needed;
  • Weakness: large magnitude ≠ importance (a small weight with large gradient can also be critical);
  • Status: the baseline and still the industry's default starting point.

2. Iterative Pruning ​

Pruning too much at once collapses accuracy — prune in multiple rounds instead:

text
Train → prune 10% → fine-tune to recover → prune 10% → fine-tune to recover → ... → reach the target sparsity

Each round of fine-tuning lets the model redistribute importance among the "remaining weights" — far better accuracy than pruning 50% in one shot.

3. Taylor Pruning (First-Order Importance) ​

Use gradient information to assess weight importance:

text
score(w) = |w · ∂L/∂w|   # weight × gradient, approximating "how much the loss rises if this is removed"

Rationale: the loss change from removing weight w ≈ w · ∂L/∂w (first-order Taylor expansion). More accurate than magnitude, but requires forward and backward passes for gradients.

4. Movement Pruning ​

Not based on the "current magnitude" but on the "trend of magnitude during training":

text
score(w) = Σ_t  w_t · ∂L/∂w_t    # accumulated: "drifting toward zero or away from it"

Drifting further from 0 → important (keep); drifting closer to 0 → unimportant (prune). Suited for pruning during fine-tuning (e.g. pruning while fine-tuning BERT).

Quick Algorithm Selection Guide

  • CV models: magnitude pruning + iterative fine-tuning — mature in practice;
  • Transformer fine-tuning: Movement Pruning (fits fine-tuning);
  • LLMs: magnitude + activation-aware (see Wanda);
  • Chasing extreme sparsity: Taylor + iterative + Lottery Ticket.

4. NVIDIA 2:4 Structured Sparsity ​

Hardware Support ​

Tensor Cores since NVIDIA Ampere (A100) support 2:4 structured sparsity (also called 1:2 sparsity): in every group of 4 consecutive weights, exactly 2 are zero — and the hardware accelerates matmul by 2×.

text
Dense weights (in groups of 4):
[1.2  0.5  0.8  0.3]   [0.4  0.9  0.1  0.7]   ...

After 2:4 sparsification (2 nonzeros per group):
[1.2  0    0.8  0  ]   [0    0.9  0    0.7]   ...
                       ↑
                Tensor Cores accelerate with dedicated instructions: matmul 2×
  • Strength: genuine 2× hardware acceleration (A100/H100 Sparse Tensor Cores);
  • Limitation: the sparsity rate is fixed at 50% (2/4 = 50%) — not tunable;
  • Tooling: NVIDIA sparsetrainer, PyTorch 2:4 sparsity support.

The Best Combo for 2:4 Sparsity

2:4 sparsity + INT4 quantization: first use magnitude pruning to reach 2:4 sparsity (~1% loss), then INT4 quantization (~1% loss) — overall that's 4× memory (quantization) + 2× compute (sparsity) = 8× total gain at ~2% accuracy loss. On H100, this is an extremely cost-effective path to extreme compression.

Other Structured Sparsity Formats ​

  • 1:2 / 1:4 / 1:8: higher sparsity but little hardware support;
  • Block sparsity: sparsity in blocks (used by the sparse variant of FlashAttention);
  • Head sparsity: prune entire attention heads — viable on LLMs but accuracy risk is high.

5. Pruning in the LLM Era ​

Traditional pruning runs into trouble on LLMs:

  1. Models are large — the "prune → fine-tune" loop is too expensive;
  2. Fine-tuning data is closed (the GPT series can't be retrained);
  3. Structured pruning destroys pretrained representations, and accuracy degrades badly.

Hence LLM-specific pruning methods:

Wanda (Pruning by Weights and Activations) ​

  • Idea: assess weight importance not just by magnitude — combine it with activation magnitude: weights with large |w| · ||x|| matter;
  • Strength: no retraining or fine-tuning needed — pure forward-pass pruning; a 70B model takes hours;
  • Sparsity: 50% unstructured with <2% accuracy loss;
  • Limitation: unstructured → real speedups require 2:4 sparse rearrangement.

LLM-Pruner ​

  • Idea: relies on structured pruning + LoRA fine-tuning for recovery;
  • Flow: find dependency groups → structured pruning → LoRA fine-tune for a few hundred steps → accuracy recovered;
  • Sparsity: 20-30% structured with <5% loss;
  • Strength: prunes whole heads/MLP dimensions — hardware-acceleratable.

ShortGPT (Layer Pruning) ​

  • Idea: remove redundant entire layers from the Transformer — some layers have outputs nearly identical to their inputs and can be skipped;
  • Sparsity: remove 25% of layers with <2% accuracy loss (validated on LLaMA-2-70B);
  • Strength: removing whole layers → shallower model → real speedup;
  • Limitation: remove the wrong layers and accuracy collapses — inter-layer similarity must be computed to pick which ones.

The Practical Value of LLM Pruning Is Limited

In the LLM era, PTQ quantization has largely replaced pruning:

  • Quantization (W4A16) compresses 4× with ~1-3% accuracy loss and no training;
  • After 50% pruning, LoRA fine-tuning is usually needed to recover accuracy — and the compression ratio may still lose to quantization. When LLM pruning still makes sense:
  • Extremely tight memory (edge deployment): quantization + sparsity stacked;
  • Extreme compute (H100 sparse Tensor Cores): 2:4 sparsity + INT4;
  • Whole-layer removal (ShortGPT): a shallower model gives real speedups. For ordinary online LLM serving, default to Weight-Only Quantization and Mixed Precision and treat pruning as a complement.

6. Sparsification and Acceleration Paths ​

Different sparsity rates call for different acceleration paths:

SparsitySparsity TypeAcceleration MethodBest For
50% (2:4)StructuredNVIDIA Sparse Tensor CoresH100/A100, 2× speedup
70-90%UnstructuredDedicated sparse kernels (e.g. DeepSparse)CPU edge deployment
90%+ExtremeSoftware simulationResearch
Whole-layer removalStructuredShallower modelAny hardware
Head pruningStructuredSmaller modelAny hardware

7. Trade-offs ​

  • Quantization vs. pruning: quantization first for LLM inference (cheaper), pruning as a complement; comparable for CV models;
  • Structured vs. unstructured: deployment favors structured (real speedups); research favors unstructured for accuracy;
  • Magnitude vs. Taylor: accuracy-first picks Taylor; simple deployment picks magnitude;
  • LLM scenarios: PTQ quantization leads; stack 2:4 sparsity on top (quantization + sparsity = 8× compression);
  • Training resources: if you can fine-tune, use iterative + LoRA recovery; if not, use Wanda (pure forward pass).

Further Reading ​

References ​