Appearance
Quantization Classics: LLM.int8 / GPTQ / AWQ
LLM quantization followed a steep curve, packing an entire era into 2022-2023: LLM.int8() first got a 175B model running on a single GPU (merely "not breaking"), GPTQ pushed 3/4-bit quantization to near-FP16 accuracy ("usable"), and AWQ used activation awareness to make 4-bit sturdier and more hardware-friendly ("pleasant to use"). This page walks through the three milestones one by one, briefly adds SmoothQuant (the W8A8 route), and closes with engineering guidance for choosing between them.
Before reading this, work through the Quantization concept page — symmetric/asymmetric quantization, scale/zero-point, and the difference between weight-only and W8A8. None of that is repeated here.
Comparison at a Glance
| Method | Bit width | Accuracy loss | Speed gain | Memory gain | Year / Venue |
|---|---|---|---|---|---|
| LLM.int8() | Weights INT8 (activations FP16) | Zero degradation at 175B | On par with FP16 (INT8 matmul + the FP16 outlier part) | Weight memory halved | 2022 / NeurIPS 2022 |
| GPTQ | Weights 3/4-bit (weight-only) | 3/4-bit close to FP16 | End-to-end: ~3.25x on A100, ~4.5x on A6000 (vs. FP16) | 3-4x weight compression | 2023 / ICLR 2023 |
| AWQ | Weights 4-bit (weight-only) | Robust, generalizes well (does not depend on calibration-set distribution) | TinyChat 4-bit is 3x+ faster than FP16 | ~4x weight compression | 2023 / MLSys 2024 (Best Paper) |
| SmoothQuant (brief) | 8-bit weights + activations (W8A8) | Negligible | 1.56x speedup | 2x memory reduction | 2023 / ICML 2023 |
What fundamentally separates the three
LLM.int8 answers "why does INT8 break?", GPTQ/AWQ answer "how do you compress weights to 4-bit without losing accuracy?", and SmoothQuant answers "how do you quantize the activations too, so you can actually collect the hardware acceleration?". They differ completely in optimization target, bit width, and use case — the question is never "which is better" but "where is your bottleneck."
1. LLM.int8(): Fitting a 175B Model on a Single GPU (Dettmers et al., NeurIPS 2022)
One-Sentence Contribution
Outliers in LLM activations are the culprit behind INT8 quantization breaking down. Split the outlier feature dimensions out to run in FP16 while the other 99.9% of values run in INT8 (mixed-precision decomposition), and a 175B model can do INT8 inference on a single GPU with zero performance degradation.
Background and Motivation
INT8-quantizing small models (≤1B) usually costs barely any accuracy — engineers have done it for years. But apply the same recipe to large models ≥6.7B and accuracy collapses dramatically. The paper traced the cause not to the weights but to the activations:
- Large models develop systematic outlier features (emergent features): on a handful of hidden dimensions the activations are far larger than on the rest (they can exceed 6 standard deviations), and every transformer layer has them.
- These outliers dominate attention and predictions. They blow out INT8's quantization range, and the accuracy of the remaining 99% of normal values is severely sacrificed as a result.
The Method: Mixed-Precision Decomposition
The idea in one line: don't let the outliers ruin everyone else — pull out the few troublemakers and compute them in FP16, and run the bulk in INT8.
text
Every linear layer (e.g., the attention Q/K/V projections, MLP):
Activations X (rows) × Weights W (columns)
│
▼
Detect outlier dimensions: max activation in a column |X_i| > threshold (e.g., 6.0)
│
┌─────┴──────────────┐
▼ ▼
Non-outlier columns (>99.9%) Outlier columns (usually a few dozen dimensions at most)
Per-row/per-column scales Kept in FP16
INT8 matmul FP16 matmul
└─────┬──────────────┘
▼
Dequantize to FP16, add = final output- Vector-wise quantization: compute a scale for each row of activations and each column of weights, instead of a single scale for the whole matrix — this is the foundation of its stability compared to naive INT8.
- Outlier detection + splitting: find the columns whose activation magnitudes exceed the threshold, split out the corresponding weight and activation columns together, and multiply them exactly in FP16; the remaining columns go through INT8. In practice over 99.9% of values still travel the INT8 path, so the tiny FP16 fraction barely slows things down.
Key Results
| Metric | Numbers |
|---|---|
| Model scale | OPT-175B / BLOOM-176B, INT8 inference on a single GPU (consumer cards like the A6000/RTX 3090) |
| Accuracy | Zero performance degradation vs. FP16 (perplexity and downstream tasks identical) |
| Memory | Weight memory halved (FP16 → INT8): 175B weights drop from ~350GB to ~175GB |
| Speed | Roughly on par with FP16 (INT8 matmul plus a little FP16 mixed in, on A100) |
| Adoption | Integrated into bitsandbytes / Hugging Face Transformers; enabled with a one-line flag |
Limitations and Follow-ups
- Only the weights are quantized; activations stay FP16 — so you never get the full speedup of INT8 kernels (it saves memory, not compute bandwidth).
- The outlier phenomenon only becomes pronounced at ≥6.7B; small models see limited or no benefit.
- Follow-ups: the same team's QLoRA (arXiv 2305.14314) combined 4-bit quantization with LoRA fine-tuning; SmoothQuant later filled in the "activations should be 8-bit too" half of the story.
Why It Matters Today
The load_in_8bit=True flag is a direct product of this paper. When you need to get a 175B model running on a single card for validation, LLM.int8 is still the fastest route. Its methodology (observe an anomaly → localize it to specific dimensions → handle those precisely) is also worth copying for other numerically sensitive engineering problems.
2. GPTQ: Quantizing 175B to 3/4-bit at Near-Lossless Accuracy (Frantar et al., ICLR 2023)
One-Sentence Contribution
Hessian-guided (second-order) layer-wise weight quantization + greedy ordering + error compensation: a 175B model is quantized to 3/4-bit in about 4 GPU-hours on a single GPU at accuracy close to FP16 — the first time a 175B model ran generative inference on one card.
Background and Motivation
After LLM.int8, the next question wrote itself: can we push down to 4-bit or below? At 4-bit, 175B weights need only ~88GB — a single A100 (80GB) plus a bit of activation memory suffices. And the smaller the weights, the less bandwidth it takes to stream them from memory during decoding, so generation gets faster.
But low-bit quantization puts extreme demands on "one-shot" quantization: you must minimize the quantization error layer by layer, without retraining via backpropagation (retraining a large model is off the table). The authors' earlier OBQ framework (optimal brain quantization) took a greedy per-weight approach with low error, but its complexity was too high to run at large-model scale.
The Method: Hessian-Based Layer-Wise Quantization
Frame "quantizing the weight matrix W of layer i" as a per-layer optimization problem: find the quantized Ŵ that minimizes ||W·x - Ŵ·x|| (using second-order activation statistics — the Hessian — to measure which weights hurt the output most when perturbed).
text
For each layer:
1. Collect activations from a small batch of calibration data; estimate the Hessian
of the layer's weights with respect to the output error
2. Greedy: quantize first the weights with the least impact on the output error
(ordered by the Hessian)
3. Error compensation: each time a weight is quantized, the error it introduces is
"propagated" back onto the remaining unquantized weights, letting them absorb it
(mathematically guaranteed not to accumulate overall)
4. Scaling trick: instead of updating weight by weight, use "lazy batch updates"
— apply the weight correction only once every K columns, cutting the complexity
from O(per-weight) down to something that can run on 175B
Output: 3-bit or 4-bit weights + per-column scales (+ a few sensitive columns kept in FP16)In essence, GPTQ is OBQ rewritten for scale: the same second-order idea, with batch updates bringing the per-element iteration cost of O(d_row × d_col²) down to something executable.
Key Results
| Metric | Numbers |
|---|---|
| Quantization time | ~4 GPU-hours on a single GPU (A100) for a 175B model |
| Accuracy | 3-bit / 4-bit close to FP16 (GPT-175B and peers, an extremely small perplexity gap); 2-bit still usable |
| End-to-end speedup | vs. FP16: ~3.25x on A100, ~4.5x on A6000 (bandwidth gains from smaller weights) |
| First ever | A 175B model doing generative inference entirely within a single GPU after quantization |
Limitations and Follow-ups
- Weight-only: activations stay FP16; only weight storage is compressed, and the activation path at inference time is unchanged.
- Depends on a calibration set: a few hundred calibration samples are needed to estimate the Hessian; accuracy drops when the calibration distribution diverges from the real one.
- Legacy: the GPTQ model format became one of the de facto standards in
llama.cpp, vLLM, and HF Transformers, and its "layer-wise + second-order information + batch updates" framework was adopted by later methods.
Why It Matters Today
GPTQ is one of the mainstream backends for 4-bit deployment today: .safetensors models exported by AutoGPTQ work out of the box in vLLM/Triton/llama.cpp. Remember its sweet spot when choosing: memory is a hard constraint, and you have a trustworthy calibration set. When your scenario calls for something less picky about data, it's AWQ's turn.
3. AWQ: Protecting Important Weights with Activation Awareness (Lin et al., MLSys 2024 Best Paper)
One-Sentence Contribution
A weight's importance is measured not by the weight itself but by activation magnitudes. Identify and protect the roughly 1% of salient channels based on activation statistics — and do it without mixed precision (apply equivalent scaling to the salient channels instead). The result: 4-bit quantization that is steadier, more hardware-friendly, and independent of backpropagation.
Background and Motivation
GPTQ leans on a calibration set plus second-order statistics, which carries an engineering pain point: results wobble whenever the calibration distribution shifts, and it never explains which weights truly cannot be touched. AWQ set out to answer a more fundamental question: is there a way to protect only the most important 1% of weights without relying on reconstruction or optimization?
The paper's observation: the weight channels you really can't touch are identified by activation magnitudes — channels with large activations correspond to weights whose quantization error does the most damage (because at inference time these weights are amplified by the activations). The weights' own numerical magnitude, by contrast, has almost nothing to do with importance.
The Method: Equivalent Scaling
The key engineering insight: protecting a salient channel does not mean keeping it in FP16 — mixed precision is expensive on hardware (it requires special kernel support). The paper proves that scaling up a salient channel's weights shrinks the quantization error proportionally: quantization error is relative, so enlarged weights face a smaller relative quantization step. Mathematically this is equivalent to protection; in implementation it is just a scaling transform applied before uniform quantization.
text
1. Calibrate: collect activation statistics from a few samples; compute per-channel
activation magnitudes
2. Identify: the channels in the top 1% of activation magnitude = salient channels
3. Scale: multiply the salient channels' weights by s > 1 and divide the
corresponding activation columns by s
(XW = X'W' is a mathematical identity; the inference output is unchanged)
s is determined via a small search (or an analytical approximation)
4. Quantize all weights uniformly to INT4/INT8 (no mixed-precision kernels)Key Results
| Metric | Numbers |
|---|---|
| Coverage | Only ~1% of weight channels protected, yet most of the quantization error is eliminated |
| Calibration cost | Activation statistics only: no backpropagation, no reconstruction loop; generalizes to code/math/multimodal without degradation |
| Accuracy | 4-bit stable for the first time on instruction-tuned models, beating contemporaneous GPTQ |
| End-to-end | TinyChat inference framework: 4-bit is 3x+ faster than HF FP16 on desktop/mobile GPUs; 70B Llama-2 deployable on a mobile GPU |
Limitations and Follow-ups
- Still a weight-only 4-bit route; activations remain FP16.
- The scaling parameter must be searched per model (the paper gives an analytical approximation, so the cost is low), but a little hyperparameter tuning room remains.
- Follow-ups: AWQ became a built-in INT4 quantization backend in vLLM and spurred exploration of 2-bit/3-bit and KV cache quantization (
KVQuant); the same team's SmoothQuant covers the W8A8 side.
Why It Matters Today
AWQ's idea of protecting a small set of channels based on the activation distribution is among the most thoroughly field-proven ideas in this space. If you are doing edge/on-device deployment (phones, laptops, Jetson), AWQ + TinyChat is the most mature 4-bit option; if you are serving LLMs on servers, one line — --quantization awq — turns it on in vLLM.
4. SmoothQuant (Brief): Making Activations 8-bit Too
LLM.int8 quantizes only the weights — why not the activations as well? Because activation outliers are far worse than weight outliers (weight distributions are relatively uniform; activations have systematic spikes). SmoothQuant's move is clever: since activations are hard to quantize and weights are easy, migrate the quantization difficulty from the activations onto the weights — divide the activations channel-wise by a scale s while multiplying the corresponding weight columns by s. It is a mathematical identity transform (XW is unchanged), but the activation distribution gets flattened enough for the whole matrix product to run in INT8.
text
Before: activations X have severe outliers -> W8A8 breaks
After the transform: X' = X / s, W' = W × s
X'W' ≡ XW (identity)
After quantization: both X' and W' fit in INT8 -> every matmul runs on INT8 kernelsThe result: W8A8 becomes viable for LLMs for the first time, with roughly 1.56x speedup and 2x memory reduction, and a 530B model can be served within a single node (ICML 2023).
Choosing between the three routes
- Want to get a 175B validation running fastest → LLM.int8 (zero setup, a one-line flag);
- Want to compress to 4-bit to save memory → GPTQ (when your data is trustworthy) or AWQ (when you want more robustness / edge targets);
- Want to benefit from INT8 kernel acceleration (TensorRT/Triton INT8 backends) → SmoothQuant W8A8. The full decision matrix is on Quantization.
5. Engineering Recommendations
| Your scenario | Recommendation | Why |
|---|---|---|
| Running 70B-175B on one card, memory is the only bottleneck | GPTQ 4-bit / AWQ 4-bit | 4x weight compression, faster end-to-end |
| Your serving stack already uses INT8 hardware kernels (TensorRT, Triton) | SmoothQuant (W8A8) | Both weights and activations at 8-bit, fully utilizing the hardware |
| Rapid experiments, no appetite for building a calibration set | LLM.int8 (bitsandbytes) | No calibration, works out of the box |
| Phone / edge / consumer GPU | AWQ (TinyChat ecosystem) | Hardware-friendly, fused 4-bit kernels |
| Model under 1B | Quantize cautiously, or prefer distillation | LLM quantization methods pay off little on small models (see below) |
6. Shared Limitations
- Calibration-set dependence: GPTQ, AWQ, and SmoothQuant all need a small batch of calibration data (usually a few hundred samples). When the calibration distribution diverges from real production traffic, accuracy drops noticeably — always calibrate and validate with real traffic samples before going live.
- Poor results on small models: LLM.int8's outlier pattern only becomes significant at 6.7B+; GPTQ/AWQ offer little benefit below 1B and can even end up worse than FP16. For small models, consider distillation first — see Distillation classics.
- Weight-only doesn't speed up compute: GPTQ/AWQ compress storage/bandwidth, but the decode-stage matmuls still run on FP16 activations — for compute speedup you need W8A8 or structured sparsity.
- Numerical safety: INT8/INT4 inference calls for validating the output distribution and monitoring for anomalous tokens; in production, keep an FP16 shadow route for comparison. See Monitoring and observability.
Further Reading
- Quantization — the conceptual foundation: scale, weight-only, W8A8, calibration
- Distillation classics: KD and the distillation family — the distill-first-then-quantize combo
- PagedAttention: the vLLM system paper — stacking quantization with KV cache optimization
- GPUs and hardware selection — how memory bandwidth determines your quantization gains
- Choosing frameworks and platforms — GPTQ/AWQ support across frameworks
References
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (arXiv 2208.07339)
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (arXiv 2210.17323)
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (arXiv 2306.00978)
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models (arXiv 2211.10438)
- QLoRA: Efficient Finetuning of Quantized LLMs (arXiv 2305.14314)