Skip to content

Quantization

At a glance Quantization represents weights and activations with lower-precision numbers—the best value-for-money acceleration and slimming technique available. This article covers the FP16/BF16/INT8/INT4 trade-offs, PTQ vs QAT, weight vs activation quantization, and GPTQ/AWQ/GGUF in the LLM era.

Quantization ​

The one-sentence definition: quantization replaces the high-precision numbers in a model (FP32/FP16) with lower-bit integers (INT8/INT4)—a compression and acceleration technique that trades a controllable amount of accuracy for outsized gains in size, VRAM, bandwidth, and compute.

Industry insight: quantization is currently the highest-value move in deployment optimization. FP32→INT8 cuts weight size by 75% and speeds up inference 2–4×, with accuracy loss under 1% on most tasks; INT4 halves it again, letting a 7B model fit in 4GB of VRAM. Meta, NVIDIA, and Microsoft all run INT8/INT4 inference at scale in their production systems (the LLM.int8 paper reports BLOOM-176B quantized to INT8 with no accuracy loss). The hard part isn't "quantization itself" but quantization discipline: who runs calibration well, who controls error well, who validates on same-distribution data—the gap is all in the second half of the process.

1. What Quantization Is and Why It's Worth Doing ​

1.1 What FP32 → INT8 Physically Saves ​

BenefitQuantizing FP32→INT8Why
Weight size75% smaller4 bytes → 1 byte
VRAM usage75% smallerSame as above; fits 4× larger models
Inference speed (when bandwidth-bound)~2× fasterNearly half the weight-read time
Compute (INT8 tensor cores)2–4×More integer multiply-adds per cycle
EnergyLowerInteger math draws less power

Why quantization helps inference more than training: inference is bandwidth-intensive (a few operations per byte of weights read), so halving the weight bytes halves the wait for data to arrive; training needs gradient precision, and quantization breaks backpropagation—the gains are far smaller than the risks (see Inference: From Forward Pass to Inference Engines).

1.2 The Math Underneath ​

Map a floating-point range onto an integer range:

text
float_value ≈ scale × (int_value - zero_point)

Symmetric quantization (the common case):
  scale = max(|float_min|, |float_max|) / 127     # INT8
  int_value = round(float_value / scale)

The error source is immediately visible: quantization squeezes a continuous interval into 256 buckets (INT8). Outliers stretch the buckets and crowd ordinary values into a handful of them, and accuracy collapses. That's the starting point for every advanced method below.

2. The Precision Format Landscape ​

FormatBitsDynamic rangePrecisionTypical use
FP3232~1e-38–3e38HighTraining, accuracy baseline
FP1616Limited (±65504)MediumDefault for general inference
BF1616Same as FP32Low (short mantissa)Training/inference, overflow-tolerant
INT88-128–127 (integers)Medium-lowInference workhorse, 75% smaller
INT44-8–7LowLLM slimming, needs group calibration
FP8 (E4M3/E5M2)8LimitedMediumNext-gen training/inference

Why Training Uses BF16, Not FP16

FP16's dynamic range is too narrow—gradients overflow/underflow during training. BF16 has FP32's dynamic range and only sacrifices precision, so training stays stable. Inference is the opposite: weight ranges are known, so FP16's extra precision wins, making it the inference default. Don't mix the two up.

3. Two Routes: PTQ and QAT ​

3.1 Post-Training Quantization (PTQ) ​

After training, run a small calibration set through the model to gather weight/activation value ranges, then convert straight to INT8. Cheap (tens of minutes), no retraining—the default route.

text
Trained model
   │ collect calibration data (a few hundred to a few thousand samples, must match the production distribution)
   │ gather per-layer/per-channel min/max (or percentiles)
   │ compute scale / zero_point
   ▼
INT8 model → compare metrics on a validation set → ship if the drop is acceptable

3.2 Quantization-Aware Training (QAT) ​

Insert "fake quantization" ops during training so the model learns to be robust to quantization error. Best results (usually <0.5% drop) but requires retraining/fine-tuning—expensive.

text
Training graph with fake-quant ops: forward simulates INT8 rounding; backward updates weights via the straight-through estimator (STE)
After fine-tuning → remove fake quant → export the INT8 model
DimensionPTQQAT
CostHoursDays (needs GPU training resources)
AccuracyManageable on large models; small/activation-sensitive models drop moreUsually best
Data needsCalibration dataTraining data + a training pipeline
WhenDefault routeWhen PTQ's drop is unacceptable

Rule of thumb: PTQ first; if the drop exceeds 1% (or the business can't accept it), go QAT or distillation (see Distillation, Pruning, and Low-Rank Factorization). LLMs have huge parameter counts and high redundancy, so PTQ usually works well.

4. What to Quantize: Weights, Activations, KV Cache ​

4.1 Weights vs Activations ​

TargetDifficultyWhyBenefit
WeightsEasyStatic range, can be profiled offline75% smaller, half the bandwidth
ActivationsHardRange varies with input, outliers aboundCompresses activations too, further cutting memory/bandwidth
KV cacheMediumIntermediate state that grows with tokensCritical for long-context/high-concurrency LLMs (see LLM Inference Optimization)

The root of the activation problem is outliers: within a layer, a few channels can be 100× larger than the rest. Set the scale by the global max and every other channel collapses into nearly a single value. That's why activation quantization is where LLMs lose the most accuracy.

4.2 Granularity: per-tensor / per-channel / per-group ​

GranularityRange sharing a scaleAccuracyStorage overhead
per-tensorOne scale per layerLowNearly zero
per-channelOne scale per output channelMediumSmall
per-group (e.g., 128 weights)One scale per groupHigh2–4 bytes per 128 weights

The usual sweet spot: per-channel or per-group for weights, per-tensor for activations (per-channel activations are hard to implement efficiently at inference time).

4.3 Taming Activation Outliers: SmoothQuant's Idea ​

SmoothQuant (ICML 2023) is an elegant move: don't touch the hard-to-quantize activations—migrate their outliers into the easy-to-quantize weights—divide the weights by a per-channel smoothing factor and multiply the activations by it. Activations become well-behaved, so per-tensor INT8 can quantize both weights and activations without losing accuracy. It's a cornerstone of today's W8A8 (both weights and activations at 8-bit) schemes.

5. Quantization in the LLM Era ​

5.1 GPTQ: Greedy Search + Hessian Approximation ​

GPTQ (2022) builds on the OBC ("optimal brain compression") idea: quantize weights column by column, using a Hessian approximation to adjust the remaining weights and compensate for the output error each quantized column introduces. Accurate with fast calibration, it's the classic 4-bit scheme, supported by vLLM, TGI, and other frameworks.

5.2 AWQ: Protecting Important Weights by Activation Magnitude ​

AWQ (2023) observed that a weight's importance is determined by the magnitude of the activations flowing through it. It profiles the activation magnitude distribution and scales up "important" weight channels to protect them—no backpropagation or retraining needed—preserving accuracy even at INT4. It's commonly paired with TensorRT-LLM or vLLM in production.

5.3 GGUF's q4_K_M and Friends ​

q4_K_M in GGUF is the llama.cpp ecosystem's block-wise quantization: every 256 weights form a block carrying its quantization type and scale, and _M means mixed precision (the more important parts get higher precision). Strengths: works out of the box, with CPU/local deployment as a first-class citizen (see Model Formats and Conversion). Its limitation: tightly coupled to the llama.cpp runtime.

MethodBit widthCalibration needsAccuracy drop (rules of thumb)Production ecosystem
LLM.int8INT8Small~0Built into Hugging Face
GPTQINT4/INT8SmallControllable, 1–3%vLLM/TGI/AutoGPTQ
AWQINT4SmallBetter than GPTQvLLM/TensorRT-LLM
GGUF q4_K_M~4.5 bitsNone (built in)2–5% (acceptable locally)llama.cpp/Ollama
SmoothQuantINT8 W8A8SmallLowestBuilt into some frameworks

6. Where Quantization Error Comes From and How to Contain It ​

Error sourceMechanismMitigation
OutliersA few huge values inflate the scale, crushing everyone else's precisionper-channel/per-group, SmoothQuant, AWQ's protection
Range mismatchThe calibration set doesn't match the production distributionCalibration data must come from the same distribution (the most common failure)
Error accumulationErrors compound layer by layerFiner granularity, mixed precision (leave some layers unquantized)
Activation quantizationActivation ranges shift with inputProfile activation ranges with a calibration set, add SmoothQuant

Calibration Data Is the Lifeline of Quantization

Calibrate on 1000 random training-set samples, then serve a production distribution that looks nothing like them, and quantization error blows up. Sample calibration data from the same distribution as production, and after quantization, validate on a same-distribution test set.

7. Acceptance Discipline After Quantization ​

  1. Run same-distribution data: before/after comparisons must use the same production-distribution data—compare output distributions layer by layer, or compare task metrics directly.
  2. Judge by business metrics, not bit equality: an AUC/accuracy/Rouge drop < 1% is usually acceptable; bit-level differences between fp16 vs int8 are inevitable—don't chase them.
  3. Cover long-tail inputs: validate abnormal inputs and extreme text separately—quantized models are more sensitive to outlier inputs.
  4. Keep monitoring after launch: a quantized model's online baselines may differ from FP16's, so give it its own baselines in monitoring (see Monitoring and Observability).

Trade-offs ​

Decision pointOptionsHow to choose
PTQ vs QATEasy vs more accurateDefault PTQ; go QAT/distillation if the drop is unacceptable
INT8 vs INT4Stable vs leanerStart with INT8 server-side; INT4 for LLMs/edge
Weights only vs weights+activationsSimple vs strongerWeights first; add activation quantization if bandwidth is still the bottleneck
per-tensor vs per-channelSimple vs more accurateper-channel weights, per-tensor activations
GPTQ vs AWQ vs GGUFEcosystem differencesAWQ/GPTQ for production services; GGUF for local tools

One-line summary: quantization is the engineering of "trading precision for resources"—pick the right route (PTQ first), manage the error (same-distribution calibration data), and hold the acceptance line (no business-metric drop), and you can improve a model's size and speed by an order of magnitude in one stroke.

Further Reading ​

References ​