Appearance
Model Quantization Fundamentals
Concept Definition: Representing the Same Information with Fewer Bits
Quantization maps high-precision floating-point numbers (FP32/FP16/BF16) to low-precision fixed-point numbers (INT8/INT4/FP8/FP4) so that inference can compute in low precision, reducing memory footprint and bandwidth pressure. It is the most cost-effective technique in LLM inference acceleration — bar none.
Two key insights for understanding quantization:
- Quantization is "lossy compression of information" — going from 16 bits to 4 bits means discarding 75% of the precision bits, so accuracy loss is inevitable; but model weights are typically concentrated in a narrow range (roughly Gaussian) with few outliers, so "compress small but stay usable" is feasible;
- The core value of quantization in LLM inference is cutting bandwidth, not compute — the decode phase is memory-bound, and INT4 weights are 4× smaller than FP16 in bytes, translating directly into 4× speedup; as for compute, Tensor Core INT8 is even faster than FP16.
The engineering value of quantization (typical numbers):
| Scheme | Memory | Speed | Accuracy Loss | Details |
|---|---|---|---|---|
| FP16/BF16 (baseline) | 1× | 1× | 0 | — |
| W8A16 (weight-only) | 2× | ~1.5× | <1% | Weight-Only Quantization and Mixed Precision |
| W4A16 (weight-only) | 4× | ~3-4× | 1-3% | Weight-Only Quantization and Mixed Precision |
| W8A8 (weights + activations) | 2× | ~2× | 1-5% | SmoothQuant |
| FP8 (H100+) | 2× | ~2× | <1% | This page |
| FP4 (B200+) | 4× | ~4× | 2-5% | This page |
1. Linear Quantization: The Most Basic Mapping
The Formula
To quantize an FP16 tensor x to INT8:
text
q = round(x / scale) + zero_point (quantize)
x' = (q - zero_point) × scale (dequantize)- scale: a floating-point scaling factor that determines "how large a floating-point range each quantization level represents";
- zero_point: the fixed-point value that corresponds to floating-point 0;
- q: the quantized integer, taking values in the range [0, 2^bit - 1].
Symmetric vs. Asymmetric Quantization
| Type | zero_point | Formula | Applies To |
|---|---|---|---|
| Symmetric quantization | =0 (or midpoint) | q = round(x / scale), range [-127, 127] | Weights (distribution is usually symmetric around 0) |
| Asymmetric quantization | ≠0 | q = round(x / scale) + zp, range [0, 255] | Activations (often offset, e.g. non-negative after ReLU) |
Why Weights Use Symmetric and Activations Use Asymmetric
Weight distributions are usually symmetric about 0 (Gaussian) — symmetric quantization saves one zero-point parameter, and dequantization needs no zp addition, making computation faster. Activation distributions are often asymmetric (all positive after ReLU, range [0,1] after Softmax) — asymmetric quantization makes fuller use of the 256 INT8 levels.
Sources of Accuracy Loss
- Quantization error: floating-point values rounded to integers;
- Clipping error: values beyond [-127, 127] are clipped;
- Outliers: a few extremely large weights stretch the scale and compress the precision of all other weights;
- Activation distribution drift: at serving time, activation ranges exceed the calibrated ranges, invalidating the scale;
- Cross-layer error accumulation: each layer's quantization error propagates backward, amplified in deep networks.
2. Quantization Granularity: How Coarse the Scale Sharing Is
At what granularity is the scale shared? This is the core knob of quantization:
| Granularity | Number of Scales | Advantages | Disadvantages |
|---|---|---|---|
| Per-tensor | 1 scale | Simplest to implement, hardware-friendly | Outliers amplify the error; poor accuracy |
| Per-channel | 1 per row/column | Good accuracy (precision-sensitive) | More scale parameters |
| Per-group | 1 per N consecutive values | Best accuracy; the standard for INT4 | Complex to implement; unpack overhead |
| Per-token | 1 per activation token | Adapts to activation distribution changes | Activations only |
| Per-head | 1 per attention head | Adapts to inter-head differences | Mostly for KV cache quantization |
INT4 weight quantization almost always uses per-group (group_size=128) — finer groups give better accuracy, but each additional scale adds 2 bytes of metadata. At group_size=128, metadata overhead is 2/128 = 1.5%, which is acceptable.
The Engineering Trade-offs Behind Mainstream Choices
- W8A16 weights: per-channel is enough (INT8 tolerance is large);
- W4A16 weights: per-group=128 (INT4 requires fine-grained groups to preserve accuracy);
- W8A8 activations: per-token (activation distributions vary greatly with input; dynamic scales needed);
- KV cache: per-token + per-head (long contexts are precision-sensitive).
3. PTQ vs. QAT: To Retrain or Not
Post-Training Quantization (PTQ)
Quantize directly after training — no retraining required. Run a forward pass over calibration data to collect per-layer activation statistics and determine the scales.
- Pros: zero training cost, done in a few hours;
- Cons: accuracy loss can be large at low bit-widths (INT4);
- Mainstream algorithms: GPTQ, AWQ, SmoothQuant, ZeroQuant.
Quantization-Aware Training (QAT)
Simulate quantization error during training (forward pass uses quantized pseudo-weights; backward pass uses the STE straight-through gradient), so the model learns to "behave well under quantization."
- Pros: better accuracy at low bit-widths (INT4 near the original precision);
- Cons: requires training data and compute — expensive;
- Mainstream algorithms: LLM-QAT, SmoothQuant-QAT.
QAT Has Faded in the LLM Era
In the classical CV era, QAT was the mainstream for INT8 quantization. But LLMs are huge, training is expensive, and training data is closed — almost all LLM quantization is PTQ. QAT is only worth considering for extreme low bit-widths (INT2, INT3) or extreme accuracy requirements, and it needs the original training data or large-scale synthetic data.
4. Calibration: The Key to Finding the Right Scale
The core step of PTQ is calibration — run representative inputs through a forward pass, collect min/max (or more sophisticated statistics) of each layer's activation distribution, and then decide the scales:
| Method | Principle | Advantages | Disadvantages |
|---|---|---|---|
| MinMax | Take the max absolute value of activations | Simple | Outliers amplify the error |
| Percentile | Take the 99.9th percentile | Outlier-resistant | Hyperparameter tuning needed |
| MSE | Minimize MSE before/after quantization | General-purpose | Slow optimization |
| ACIQ | Assume Gaussian/Laplace distribution and compute the optimal threshold analytically | Closed-form solution | Assumptions may not hold |
| ZNWC | Zero-point + normalized weighted clustering | Adapts to multi-modal distributions | Complex to implement |
Calibration data must be representative — a sample set of typical production prompts, usually 500-2000 examples. Calibration set bias (e.g. calibrating a Chinese model on pure English data) is one of the leading causes of PTQ accuracy blowups.
5. Mainstream Algorithm Families: The "Big Four" of LLM Quantization
GPTQ (Generative Pre-trained Transformer Quantization)
- Idea: second-order quantization error minimization based on the Hessian matrix;
- Key contribution: quantize weights column by column, and after each column use the Hessian to adjust the remaining columns to compensate for the error;
- Strength: excellent INT4 accuracy (<1% loss);
- Engineering: AutoGPTQ; vLLM ships a built-in GPTQ Marlin kernel;
- Details: Weight-Only Quantization and Mixed Precision.
AWQ (Activation-aware Weight Quantization)
- Idea: protect the salient weights — channels with large activation magnitudes correspond to more important weights, so keep high precision for them during quantization;
- Key trick: equivalent scaling (W' = W/s, x' = s·x) shifts the salient weights into a quantization-friendlier region;
- Strength: INT4 accuracy matches or beats GPTQ, and faster than GPTQ (more efficient kernels);
- Engineering: vLLM's AWQ kernel, TensorRT-LLM AWQ;
- Details: Weight-Only Quantization and Mixed Precision.
SmoothQuant
- Idea: solve the activation outlier problem — a few extreme values in LLM activations wreck INT8 activation quantization;
- Key trick: equivalently shift the "difficulty" from activations onto weights (W' = W/s, x' = s·x smooths the activations and steepens the weights — but weights are easy to quantize anyway);
- Strength: makes W8A8 quantization on LLMs viable for the first time;
- Use case: the W8A8 mode (both weights and activations in INT8), suited for compute-intensive prefill.
The ZeroQuant Family (from DeepSpeed)
- Idea: layer-wise PTQ + per-token activation quantization;
- ZeroQuant-Z: zero-point symmetric quantization for weights;
- ZeroQuant-Infinity: scales to trillion-parameter models;
- Use case: integrated in the DeepSpeed ecosystem, but its community is smaller than GPTQ/AWQ.
6. FP8 and FP4: The New Direction in the H100/B200 Era
FP8 (Standard on H100/H200/H20)
- Two formats: E4M3 (4 exponent bits + 3 mantissa bits, higher precision) and E5M2 (5 exponent bits + 2 mantissa bits, larger dynamic range);
- Native Tensor Core support → inference speed close to INT8 with far better accuracy;
- No calibration needed (FP8 carries its own dynamic range);
- Implemented in NVIDIA Transformer Engine;
- Best for: H100+ GPUs and scenarios that demand lossless-ish accuracy.
FP4 (Standard on Blackwell B200/B100)
- 4-bit floating point, paired with microscaling (a group of 16 values sharing one scale);
- Compute ceiling of 2.25 PFLOPS (FP4), ~4× that of FP16;
- Still requires PTQ calibration (choosing the microscaling scales);
- Best for: B200 + extreme-throughput scenarios.
Why FP8/FP4 Are the Future
- Floating-point formats are more accurate than integer formats — the exponent bits give a built-in dynamic range, so no per-group scales are needed;
- Native Tensor Core acceleration — no unpacking required (INT4 must be unpacked to INT8 before compute);
- Lighter calibration — basically just pick the format and the microscaling granularity;
- Limitation: requires the new-generation hardware (FP8 from H100, FP4 from B200) — it won't run on A100.
7. Quantization Accuracy Loss: Common Failure Modes
| Problem | Symptom | Cause | Countermeasure |
|---|---|---|---|
| Activation outliers | INT8 activation quantization collapses | Extreme values in LLM attention outputs | Use SmoothQuant to shift the difficulty onto weights |
| Low-bit weights | Poor INT4 accuracy | Outlier weights dominate the scale | Use AWQ to protect salient weights |
| KV cache quantization | Accuracy degrades at long context | Accumulated KV error | Per-token per-head quantization |
| Calibration set shift | Good offline, poor online | Calibration data distribution doesn't match production | Calibrate with typical production samples |
| Missing mixed-precision layers | A few layers collapse | A global quantization policy is unkind to sensitive layers | Keep critical layers (softmax, final norm) in FP16 |
Don't Blindly Quantize Everything to INT4
- W4A16 (INT4 weights + FP16 activations): the best value in the decode phase — the current default choice for LLM quantization;
- W4A4: extreme compute efficiency but heavy accuracy loss — recommended only on B200+;
- W8A8: good accuracy, but saves no decode bandwidth (INT8 weights are still a 2× compression, and compressing activations doesn't reduce weight traffic);
- Mixed precision: keep sensitive layers such as attention softmax and final norm in FP16, quantize the rest — nearly lossless.
8. Trade-offs
- PTQ vs. QAT: 99% of LLM scenarios use PTQ (GPTQ/AWQ); reserve QAT for extreme low bit-widths;
- INT8 vs. FP8: choose FP8 on H100+ (better accuracy, no calibration); choose INT8 on A100 (AWQ/GPTQ);
- W4A16 vs. W8A8: decode-heavy (online serving) → W4A16; prefill-heavy (offline) → W8A8;
- Quantization vs. distillation: quantization keeps the model architecture and is cheap; distillation swaps in a smaller model and can go further but costs training — usually quantize first, then consider distillation;
- Quantization vs. pruning: quantization cuts bytes; pruning cuts the parameter count — the two stack (sparse INT4 weights).
Further Reading
- Weight-Only Quantization and Mixed Precision — engineering details of GPTQ/AWQ
- Kernel Fusion and Custom Kernels — the Marlin kernel fusing W4A16 dequantization + matmul
- The GPU Memory Hierarchy and the Bandwidth Wall — why cutting bytes translates directly into speed
- Pruning and Sparsification — the other direction of compression
- Knowledge Distillation — the "training-side" tool of compression
- vLLM and PagedAttention — the industrial implementation of AWQ/GPTQ in vLLM
- TensorRT and GPU Inference — FP8/INT8 in the NVIDIA stack
- Classic Papers in Depth — links to the original GPTQ, AWQ, and SmoothQuant papers
References
- Frantar et al. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (ICLR 2023) — the original GPTQ paper
- Lin et al. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024) — the original AWQ paper
- Xiao et al. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models (ICML 2023) — the original SmoothQuant paper
- Liu et al. ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale Transformers (NeurIPS 2022) — the original ZeroQuant paper
- NVIDIA Transformer Engine Documentation — FP8/FP4 engineering implementation
- Jacob et al. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (CVPR 2018) — the foundational paper of classical CV quantization (QAT)