Appearance
Quantization
The one-sentence definition: quantization replaces the high-precision numbers in a model (FP32/FP16) with lower-bit integers (INT8/INT4)—a compression and acceleration technique that trades a controllable amount of accuracy for outsized gains in size, VRAM, bandwidth, and compute.
Industry insight: quantization is currently the highest-value move in deployment optimization. FP32→INT8 cuts weight size by 75% and speeds up inference 2–4×, with accuracy loss under 1% on most tasks; INT4 halves it again, letting a 7B model fit in 4GB of VRAM. Meta, NVIDIA, and Microsoft all run INT8/INT4 inference at scale in their production systems (the LLM.int8 paper reports BLOOM-176B quantized to INT8 with no accuracy loss). The hard part isn't "quantization itself" but quantization discipline: who runs calibration well, who controls error well, who validates on same-distribution data—the gap is all in the second half of the process.
1. What Quantization Is and Why It's Worth Doing
1.1 What FP32 → INT8 Physically Saves
| Benefit | Quantizing FP32→INT8 | Why |
|---|---|---|
| Weight size | 75% smaller | 4 bytes → 1 byte |
| VRAM usage | 75% smaller | Same as above; fits 4× larger models |
| Inference speed (when bandwidth-bound) | ~2× faster | Nearly half the weight-read time |
| Compute (INT8 tensor cores) | 2–4× | More integer multiply-adds per cycle |
| Energy | Lower | Integer math draws less power |
Why quantization helps inference more than training: inference is bandwidth-intensive (a few operations per byte of weights read), so halving the weight bytes halves the wait for data to arrive; training needs gradient precision, and quantization breaks backpropagation—the gains are far smaller than the risks (see Inference: From Forward Pass to Inference Engines).
1.2 The Math Underneath
Map a floating-point range onto an integer range:
text
float_value ≈ scale × (int_value - zero_point)
Symmetric quantization (the common case):
scale = max(|float_min|, |float_max|) / 127 # INT8
int_value = round(float_value / scale)The error source is immediately visible: quantization squeezes a continuous interval into 256 buckets (INT8). Outliers stretch the buckets and crowd ordinary values into a handful of them, and accuracy collapses. That's the starting point for every advanced method below.
2. The Precision Format Landscape
| Format | Bits | Dynamic range | Precision | Typical use |
|---|---|---|---|---|
| FP32 | 32 | ~1e-38–3e38 | High | Training, accuracy baseline |
| FP16 | 16 | Limited (±65504) | Medium | Default for general inference |
| BF16 | 16 | Same as FP32 | Low (short mantissa) | Training/inference, overflow-tolerant |
| INT8 | 8 | -128–127 (integers) | Medium-low | Inference workhorse, 75% smaller |
| INT4 | 4 | -8–7 | Low | LLM slimming, needs group calibration |
| FP8 (E4M3/E5M2) | 8 | Limited | Medium | Next-gen training/inference |
Why Training Uses BF16, Not FP16
FP16's dynamic range is too narrow—gradients overflow/underflow during training. BF16 has FP32's dynamic range and only sacrifices precision, so training stays stable. Inference is the opposite: weight ranges are known, so FP16's extra precision wins, making it the inference default. Don't mix the two up.
3. Two Routes: PTQ and QAT
3.1 Post-Training Quantization (PTQ)
After training, run a small calibration set through the model to gather weight/activation value ranges, then convert straight to INT8. Cheap (tens of minutes), no retraining—the default route.
text
Trained model
│ collect calibration data (a few hundred to a few thousand samples, must match the production distribution)
│ gather per-layer/per-channel min/max (or percentiles)
│ compute scale / zero_point
▼
INT8 model → compare metrics on a validation set → ship if the drop is acceptable3.2 Quantization-Aware Training (QAT)
Insert "fake quantization" ops during training so the model learns to be robust to quantization error. Best results (usually <0.5% drop) but requires retraining/fine-tuning—expensive.
text
Training graph with fake-quant ops: forward simulates INT8 rounding; backward updates weights via the straight-through estimator (STE)
After fine-tuning → remove fake quant → export the INT8 model| Dimension | PTQ | QAT |
|---|---|---|
| Cost | Hours | Days (needs GPU training resources) |
| Accuracy | Manageable on large models; small/activation-sensitive models drop more | Usually best |
| Data needs | Calibration data | Training data + a training pipeline |
| When | Default route | When PTQ's drop is unacceptable |
Rule of thumb: PTQ first; if the drop exceeds 1% (or the business can't accept it), go QAT or distillation (see Distillation, Pruning, and Low-Rank Factorization). LLMs have huge parameter counts and high redundancy, so PTQ usually works well.
4. What to Quantize: Weights, Activations, KV Cache
4.1 Weights vs Activations
| Target | Difficulty | Why | Benefit |
|---|---|---|---|
| Weights | Easy | Static range, can be profiled offline | 75% smaller, half the bandwidth |
| Activations | Hard | Range varies with input, outliers abound | Compresses activations too, further cutting memory/bandwidth |
| KV cache | Medium | Intermediate state that grows with tokens | Critical for long-context/high-concurrency LLMs (see LLM Inference Optimization) |
The root of the activation problem is outliers: within a layer, a few channels can be 100× larger than the rest. Set the scale by the global max and every other channel collapses into nearly a single value. That's why activation quantization is where LLMs lose the most accuracy.
4.2 Granularity: per-tensor / per-channel / per-group
| Granularity | Range sharing a scale | Accuracy | Storage overhead |
|---|---|---|---|
| per-tensor | One scale per layer | Low | Nearly zero |
| per-channel | One scale per output channel | Medium | Small |
| per-group (e.g., 128 weights) | One scale per group | High | 2–4 bytes per 128 weights |
The usual sweet spot: per-channel or per-group for weights, per-tensor for activations (per-channel activations are hard to implement efficiently at inference time).
4.3 Taming Activation Outliers: SmoothQuant's Idea
SmoothQuant (ICML 2023) is an elegant move: don't touch the hard-to-quantize activations—migrate their outliers into the easy-to-quantize weights—divide the weights by a per-channel smoothing factor and multiply the activations by it. Activations become well-behaved, so per-tensor INT8 can quantize both weights and activations without losing accuracy. It's a cornerstone of today's W8A8 (both weights and activations at 8-bit) schemes.
5. Quantization in the LLM Era
5.1 GPTQ: Greedy Search + Hessian Approximation
GPTQ (2022) builds on the OBC ("optimal brain compression") idea: quantize weights column by column, using a Hessian approximation to adjust the remaining weights and compensate for the output error each quantized column introduces. Accurate with fast calibration, it's the classic 4-bit scheme, supported by vLLM, TGI, and other frameworks.
5.2 AWQ: Protecting Important Weights by Activation Magnitude
AWQ (2023) observed that a weight's importance is determined by the magnitude of the activations flowing through it. It profiles the activation magnitude distribution and scales up "important" weight channels to protect them—no backpropagation or retraining needed—preserving accuracy even at INT4. It's commonly paired with TensorRT-LLM or vLLM in production.
5.3 GGUF's q4_K_M and Friends
q4_K_M in GGUF is the llama.cpp ecosystem's block-wise quantization: every 256 weights form a block carrying its quantization type and scale, and _M means mixed precision (the more important parts get higher precision). Strengths: works out of the box, with CPU/local deployment as a first-class citizen (see Model Formats and Conversion). Its limitation: tightly coupled to the llama.cpp runtime.
| Method | Bit width | Calibration needs | Accuracy drop (rules of thumb) | Production ecosystem |
|---|---|---|---|---|
| LLM.int8 | INT8 | Small | ~0 | Built into Hugging Face |
| GPTQ | INT4/INT8 | Small | Controllable, 1–3% | vLLM/TGI/AutoGPTQ |
| AWQ | INT4 | Small | Better than GPTQ | vLLM/TensorRT-LLM |
| GGUF q4_K_M | ~4.5 bits | None (built in) | 2–5% (acceptable locally) | llama.cpp/Ollama |
| SmoothQuant | INT8 W8A8 | Small | Lowest | Built into some frameworks |
6. Where Quantization Error Comes From and How to Contain It
| Error source | Mechanism | Mitigation |
|---|---|---|
| Outliers | A few huge values inflate the scale, crushing everyone else's precision | per-channel/per-group, SmoothQuant, AWQ's protection |
| Range mismatch | The calibration set doesn't match the production distribution | Calibration data must come from the same distribution (the most common failure) |
| Error accumulation | Errors compound layer by layer | Finer granularity, mixed precision (leave some layers unquantized) |
| Activation quantization | Activation ranges shift with input | Profile activation ranges with a calibration set, add SmoothQuant |
Calibration Data Is the Lifeline of Quantization
Calibrate on 1000 random training-set samples, then serve a production distribution that looks nothing like them, and quantization error blows up. Sample calibration data from the same distribution as production, and after quantization, validate on a same-distribution test set.
7. Acceptance Discipline After Quantization
- Run same-distribution data: before/after comparisons must use the same production-distribution data—compare output distributions layer by layer, or compare task metrics directly.
- Judge by business metrics, not bit equality: an AUC/accuracy/Rouge drop < 1% is usually acceptable; bit-level differences between
fp16 vs int8are inevitable—don't chase them. - Cover long-tail inputs: validate abnormal inputs and extreme text separately—quantized models are more sensitive to outlier inputs.
- Keep monitoring after launch: a quantized model's online baselines may differ from FP16's, so give it its own baselines in monitoring (see Monitoring and Observability).
Trade-offs
| Decision point | Options | How to choose |
|---|---|---|
| PTQ vs QAT | Easy vs more accurate | Default PTQ; go QAT/distillation if the drop is unacceptable |
| INT8 vs INT4 | Stable vs leaner | Start with INT8 server-side; INT4 for LLMs/edge |
| Weights only vs weights+activations | Simple vs stronger | Weights first; add activation quantization if bandwidth is still the bottleneck |
| per-tensor vs per-channel | Simple vs more accurate | per-channel weights, per-tensor activations |
| GPTQ vs AWQ vs GGUF | Ecosystem differences | AWQ/GPTQ for production services; GGUF for local tools |
One-line summary: quantization is the engineering of "trading precision for resources"—pick the right route (PTQ first), manage the error (same-distribution calibration data), and hold the acceptance line (no business-metric drop), and you can improve a model's size and speed by an order of magnitude in one stroke.
Further Reading
- Distillation, Pruning, and Low-Rank Factorization — three more slimming routes beyond quantization, and how to combine them
- GPUs and Hardware Selection — how INT8's compute/bandwidth gains land on real hardware
- LLM Inference Optimization — KV cache quantization and quantization practice in LLM deployment
- Papers: Quantization Classics — LLM.int8 / GPTQ / AWQ, paper by paper
- Model Optimization in Practice — the full path from quantization experiments to production
References
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (Dettmers et al.)
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (Frantar et al.)
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (Lin et al.)
- SmoothQuant: Accurate and Efficient Post-Training Quantization for LLMs (Xiao et al.)
- NVIDIA INT8 inference and calibration whitepaper