Skip to content

Model Quantization Fundamentals

At a glance Quantization is the core technique that compresses FP16/BF16 weights and activations down to INT8/INT4/FP8/FP4 — trading a small amount of accuracy for 2-8× inference speedups and memory savings. This article covers the linear quantization formula, quantization granularity, PTQ vs. QAT, calibration methods, and the mainstream algorithm families GPTQ/AWQ/SmoothQuant/FP8/FP4.

Model Quantization Fundamentals ​

Concept Definition: Representing the Same Information with Fewer Bits ​

Quantization maps high-precision floating-point numbers (FP32/FP16/BF16) to low-precision fixed-point numbers (INT8/INT4/FP8/FP4) so that inference can compute in low precision, reducing memory footprint and bandwidth pressure. It is the most cost-effective technique in LLM inference acceleration — bar none.

Two key insights for understanding quantization:

  1. Quantization is "lossy compression of information" — going from 16 bits to 4 bits means discarding 75% of the precision bits, so accuracy loss is inevitable; but model weights are typically concentrated in a narrow range (roughly Gaussian) with few outliers, so "compress small but stay usable" is feasible;
  2. The core value of quantization in LLM inference is cutting bandwidth, not compute — the decode phase is memory-bound, and INT4 weights are 4× smaller than FP16 in bytes, translating directly into 4× speedup; as for compute, Tensor Core INT8 is even faster than FP16.

The engineering value of quantization (typical numbers):

SchemeMemorySpeedAccuracy LossDetails
FP16/BF16 (baseline)1×1×0—
W8A16 (weight-only)2×~1.5×<1%Weight-Only Quantization and Mixed Precision
W4A16 (weight-only)4×~3-4×1-3%Weight-Only Quantization and Mixed Precision
W8A8 (weights + activations)2×~2×1-5%SmoothQuant
FP8 (H100+)2×~2×<1%This page
FP4 (B200+)4×~4×2-5%This page

1. Linear Quantization: The Most Basic Mapping ​

The Formula ​

To quantize an FP16 tensor x to INT8:

text
q  = round(x / scale) + zero_point     (quantize)
x' = (q - zero_point) × scale          (dequantize)
  • scale: a floating-point scaling factor that determines "how large a floating-point range each quantization level represents";
  • zero_point: the fixed-point value that corresponds to floating-point 0;
  • q: the quantized integer, taking values in the range [0, 2^bit - 1].

Symmetric vs. Asymmetric Quantization ​

Typezero_pointFormulaApplies To
Symmetric quantization=0 (or midpoint)q = round(x / scale), range [-127, 127]Weights (distribution is usually symmetric around 0)
Asymmetric quantization≠0q = round(x / scale) + zp, range [0, 255]Activations (often offset, e.g. non-negative after ReLU)

Why Weights Use Symmetric and Activations Use Asymmetric

Weight distributions are usually symmetric about 0 (Gaussian) — symmetric quantization saves one zero-point parameter, and dequantization needs no zp addition, making computation faster. Activation distributions are often asymmetric (all positive after ReLU, range [0,1] after Softmax) — asymmetric quantization makes fuller use of the 256 INT8 levels.

Sources of Accuracy Loss ​

  1. Quantization error: floating-point values rounded to integers;
  2. Clipping error: values beyond [-127, 127] are clipped;
  3. Outliers: a few extremely large weights stretch the scale and compress the precision of all other weights;
  4. Activation distribution drift: at serving time, activation ranges exceed the calibrated ranges, invalidating the scale;
  5. Cross-layer error accumulation: each layer's quantization error propagates backward, amplified in deep networks.

2. Quantization Granularity: How Coarse the Scale Sharing Is ​

At what granularity is the scale shared? This is the core knob of quantization:

GranularityNumber of ScalesAdvantagesDisadvantages
Per-tensor1 scaleSimplest to implement, hardware-friendlyOutliers amplify the error; poor accuracy
Per-channel1 per row/columnGood accuracy (precision-sensitive)More scale parameters
Per-group1 per N consecutive valuesBest accuracy; the standard for INT4Complex to implement; unpack overhead
Per-token1 per activation tokenAdapts to activation distribution changesActivations only
Per-head1 per attention headAdapts to inter-head differencesMostly for KV cache quantization

INT4 weight quantization almost always uses per-group (group_size=128) — finer groups give better accuracy, but each additional scale adds 2 bytes of metadata. At group_size=128, metadata overhead is 2/128 = 1.5%, which is acceptable.

The Engineering Trade-offs Behind Mainstream Choices

  • W8A16 weights: per-channel is enough (INT8 tolerance is large);
  • W4A16 weights: per-group=128 (INT4 requires fine-grained groups to preserve accuracy);
  • W8A8 activations: per-token (activation distributions vary greatly with input; dynamic scales needed);
  • KV cache: per-token + per-head (long contexts are precision-sensitive).

3. PTQ vs. QAT: To Retrain or Not ​

Post-Training Quantization (PTQ) ​

Quantize directly after training — no retraining required. Run a forward pass over calibration data to collect per-layer activation statistics and determine the scales.

  • Pros: zero training cost, done in a few hours;
  • Cons: accuracy loss can be large at low bit-widths (INT4);
  • Mainstream algorithms: GPTQ, AWQ, SmoothQuant, ZeroQuant.

Quantization-Aware Training (QAT) ​

Simulate quantization error during training (forward pass uses quantized pseudo-weights; backward pass uses the STE straight-through gradient), so the model learns to "behave well under quantization."

  • Pros: better accuracy at low bit-widths (INT4 near the original precision);
  • Cons: requires training data and compute — expensive;
  • Mainstream algorithms: LLM-QAT, SmoothQuant-QAT.

QAT Has Faded in the LLM Era

In the classical CV era, QAT was the mainstream for INT8 quantization. But LLMs are huge, training is expensive, and training data is closed — almost all LLM quantization is PTQ. QAT is only worth considering for extreme low bit-widths (INT2, INT3) or extreme accuracy requirements, and it needs the original training data or large-scale synthetic data.

4. Calibration: The Key to Finding the Right Scale ​

The core step of PTQ is calibration — run representative inputs through a forward pass, collect min/max (or more sophisticated statistics) of each layer's activation distribution, and then decide the scales:

MethodPrincipleAdvantagesDisadvantages
MinMaxTake the max absolute value of activationsSimpleOutliers amplify the error
PercentileTake the 99.9th percentileOutlier-resistantHyperparameter tuning needed
MSEMinimize MSE before/after quantizationGeneral-purposeSlow optimization
ACIQAssume Gaussian/Laplace distribution and compute the optimal threshold analyticallyClosed-form solutionAssumptions may not hold
ZNWCZero-point + normalized weighted clusteringAdapts to multi-modal distributionsComplex to implement

Calibration data must be representative — a sample set of typical production prompts, usually 500-2000 examples. Calibration set bias (e.g. calibrating a Chinese model on pure English data) is one of the leading causes of PTQ accuracy blowups.

5. Mainstream Algorithm Families: The "Big Four" of LLM Quantization ​

GPTQ (Generative Pre-trained Transformer Quantization) ​

  • Idea: second-order quantization error minimization based on the Hessian matrix;
  • Key contribution: quantize weights column by column, and after each column use the Hessian to adjust the remaining columns to compensate for the error;
  • Strength: excellent INT4 accuracy (<1% loss);
  • Engineering: AutoGPTQ; vLLM ships a built-in GPTQ Marlin kernel;
  • Details: Weight-Only Quantization and Mixed Precision.

AWQ (Activation-aware Weight Quantization) ​

  • Idea: protect the salient weights — channels with large activation magnitudes correspond to more important weights, so keep high precision for them during quantization;
  • Key trick: equivalent scaling (W' = W/s, x' = s·x) shifts the salient weights into a quantization-friendlier region;
  • Strength: INT4 accuracy matches or beats GPTQ, and faster than GPTQ (more efficient kernels);
  • Engineering: vLLM's AWQ kernel, TensorRT-LLM AWQ;
  • Details: Weight-Only Quantization and Mixed Precision.

SmoothQuant ​

  • Idea: solve the activation outlier problem — a few extreme values in LLM activations wreck INT8 activation quantization;
  • Key trick: equivalently shift the "difficulty" from activations onto weights (W' = W/s, x' = s·x smooths the activations and steepens the weights — but weights are easy to quantize anyway);
  • Strength: makes W8A8 quantization on LLMs viable for the first time;
  • Use case: the W8A8 mode (both weights and activations in INT8), suited for compute-intensive prefill.

The ZeroQuant Family (from DeepSpeed) ​

  • Idea: layer-wise PTQ + per-token activation quantization;
  • ZeroQuant-Z: zero-point symmetric quantization for weights;
  • ZeroQuant-Infinity: scales to trillion-parameter models;
  • Use case: integrated in the DeepSpeed ecosystem, but its community is smaller than GPTQ/AWQ.

6. FP8 and FP4: The New Direction in the H100/B200 Era ​

FP8 (Standard on H100/H200/H20) ​

  • Two formats: E4M3 (4 exponent bits + 3 mantissa bits, higher precision) and E5M2 (5 exponent bits + 2 mantissa bits, larger dynamic range);
  • Native Tensor Core support → inference speed close to INT8 with far better accuracy;
  • No calibration needed (FP8 carries its own dynamic range);
  • Implemented in NVIDIA Transformer Engine;
  • Best for: H100+ GPUs and scenarios that demand lossless-ish accuracy.

FP4 (Standard on Blackwell B200/B100) ​

  • 4-bit floating point, paired with microscaling (a group of 16 values sharing one scale);
  • Compute ceiling of 2.25 PFLOPS (FP4), ~4× that of FP16;
  • Still requires PTQ calibration (choosing the microscaling scales);
  • Best for: B200 + extreme-throughput scenarios.

Why FP8/FP4 Are the Future

  • Floating-point formats are more accurate than integer formats — the exponent bits give a built-in dynamic range, so no per-group scales are needed;
  • Native Tensor Core acceleration — no unpacking required (INT4 must be unpacked to INT8 before compute);
  • Lighter calibration — basically just pick the format and the microscaling granularity;
  • Limitation: requires the new-generation hardware (FP8 from H100, FP4 from B200) — it won't run on A100.

7. Quantization Accuracy Loss: Common Failure Modes ​

ProblemSymptomCauseCountermeasure
Activation outliersINT8 activation quantization collapsesExtreme values in LLM attention outputsUse SmoothQuant to shift the difficulty onto weights
Low-bit weightsPoor INT4 accuracyOutlier weights dominate the scaleUse AWQ to protect salient weights
KV cache quantizationAccuracy degrades at long contextAccumulated KV errorPer-token per-head quantization
Calibration set shiftGood offline, poor onlineCalibration data distribution doesn't match productionCalibrate with typical production samples
Missing mixed-precision layersA few layers collapseA global quantization policy is unkind to sensitive layersKeep critical layers (softmax, final norm) in FP16

Don't Blindly Quantize Everything to INT4

  • W4A16 (INT4 weights + FP16 activations): the best value in the decode phase — the current default choice for LLM quantization;
  • W4A4: extreme compute efficiency but heavy accuracy loss — recommended only on B200+;
  • W8A8: good accuracy, but saves no decode bandwidth (INT8 weights are still a 2× compression, and compressing activations doesn't reduce weight traffic);
  • Mixed precision: keep sensitive layers such as attention softmax and final norm in FP16, quantize the rest — nearly lossless.

8. Trade-offs ​

  • PTQ vs. QAT: 99% of LLM scenarios use PTQ (GPTQ/AWQ); reserve QAT for extreme low bit-widths;
  • INT8 vs. FP8: choose FP8 on H100+ (better accuracy, no calibration); choose INT8 on A100 (AWQ/GPTQ);
  • W4A16 vs. W8A8: decode-heavy (online serving) → W4A16; prefill-heavy (offline) → W8A8;
  • Quantization vs. distillation: quantization keeps the model architecture and is cheap; distillation swaps in a smaller model and can go further but costs training — usually quantize first, then consider distillation;
  • Quantization vs. pruning: quantization cuts bytes; pruning cuts the parameter count — the two stack (sparse INT4 weights).

Further Reading ​

References ​