Appearance
Weight-Only Quantization and Mixed Precision
Concept Definition: The Strongest Weapon Against the Decode-Phase Bandwidth Wall
LLM inference is memory-bound in the decode phase — every generated token requires reading the entire model weights from HBM once. Weight-only quantization is purpose-built for this: compress the weights to INT4/INT8 (4×/2× fewer bytes) while keeping activations in FP16/BF16 (no accuracy loss), and decode throughput gains 3-4×.
Two key insights for understanding weight-only quantization:
- It compresses only weights, not activations — because activations are the small share (a forward pass computes only a few KB for the current token), while weights are the big share (140GB for a 70B model);
- It targets decode; prefill benefits little — prefill is compute-bound, so cutting bytes doesn't directly translate into speed; decode is memory-bound, so cutting bytes does.
Weight-only quantization is the de facto standard of LLM inference today — vLLM, TensorRT-LLM, and llama.cpp all ship built-in support.
1. The Basic Recipe of Weight-Only Quantization
text
Before: W (FP16) × X (FP16) → Y (FP16) # decode-phase bandwidth bottleneck
Quant: W (INT4) → dequant → W' (FP16) × X (FP16) → Y
↑
dequantization done on the fly before the matmulThe key question: where does dequantization happen? Two mainstream approaches:
| Approach | Flow | Advantages | Disadvantages |
|---|---|---|---|
| Offline dequantization | W INT4 → W FP16 (one-shot dequant before inference) | Simple to implement | No memory saved, no speedup — an anti-pattern |
| Online dequantization + fused kernel | Dequantize each weight group inside the matmul | 4× memory savings, 4× bandwidth savings | Kernels must be hand-optimized (Marlin) |
Don't "Dequantize Offline"
"Dequantize INT4 weights back to FP16 before inference" is the most common beginner trap — it saves no memory (you still store FP16) and is actually slower (an extra dequantization step). True weight-only quantization requires a fused kernel — dequantization happens inside the matmul: each INT4 group is dequantized and consumed immediately, and the weights stay in INT4 the whole time. The Marlin kernel exists precisely for this.
2. GPTQ: Hessian-Based Second-Order Compensation
Core Idea
GPTQ (Frantar et al., 2023) is one of the most accurate algorithms for INT4 weight quantization. Its core: quantize weights column by column, and after each column use Hessian information to adjust the remaining columns and compensate for the error.
The math:
text
Quantization error minimization: minimize || W × X - Q × X ||²
where X is the calibration data, and X × X^T ≈ H (the Hessian matrix)GPTQ uses a greedy algorithm based on second-order information (the Hessian): quantize column by column, and after each column project the error onto the remaining columns weighted by the Hessian, letting subsequent columns compensate for the error already incurred. This "compensating" quantization minimizes the accumulated error.
Engineering
python
# Simplified AutoGPTQ workflow
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
quant_config = BaseQuantizeConfig(
bits=4,
group_size=128,
desc_act=False, # whether to order columns by the Hessian
)
model = AutoGPTQForCausalLM.from_pretrained("llama-2-70b")
model.quantize(calibration_data, quant_config)
model.save_quantized("llama-2-70b-gptq-4bit")Strengths and Limitations
- Strengths: excellent INT4 accuracy (<1% loss), broad compatibility (almost every LLM has a GPTQ build);
- Limitations: slow quantization (calibration + Hessian computation takes hours for a 70B model), and kernel speed slightly trails AWQ.
3. AWQ: Protecting Salient Weights
Core Idea
The core observation of AWQ (Lin et al., 2023): about 1% of the weights in an LLM are "salient" and decisive for accuracy — typically the weights on channels with large activation magnitudes. If this 1% is protected, the accuracy loss from quantizing the other 99% is negligible.
But how should "protecting 1% of weights" be implemented? AWQ uses a clever trick: equivalent scaling.
text
Original: Y = W × X
Equivalent: Y = (W / s) × (s · X) = W' × X'
↑ ↑
scaled weights scaled activationsBy multiplying channels with large activation magnitudes by s > 1 (making the salient weights relatively smaller, landing them in an INT4-friendlier region), the corresponding channels of X are equivalently divided by s — mathematically equivalent, but the salient weights lose less during quantization.
Why AWQ Is Faster Than GPTQ
AWQ's quantization process is a lightweight search that needs no Hessian — it only observes activation magnitudes over the weights, and a few hundred forward passes determine the scale for every group. The kernel implementation is also more efficient — the AWQ kernel (vLLM implementation) typically leads the GPTQ Marlin by 10-30% at comparable accuracy.
Engineering
python
# Simplified AutoAWQ workflow
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model = AutoAWQForCausalLM.from_pretrained("llama-2-70b")
tokenizer = AutoTokenizer.from_pretrained("llama-2-70b")
quant_config = {"w_bit": 4, "q_group_size": 128, "zero_point": True}
model.quantize(tokenizer, quant_config)
model.save_quantized("llama-2-70b-awq-4bit")4. GPTQ vs. AWQ: A Practical Comparison
| Dimension | GPTQ | AWQ |
|---|---|---|
| Core algorithm | Hessian-based second-order compensation | Salient-weight protection + equivalent scaling |
| Quantization time | Slow (hours for 70B) | Fast (30-60 minutes) |
| INT4 accuracy | <1% loss | <1% loss (comparable) |
| Kernel speed | Near AWQ after Marlin optimization | Faster out of the box |
| Ecosystem | AutoGPTQ, vLLM, TensorRT-LLM | AutoAWQ, vLLM, TensorRT-LLM |
| Fine-tuning friendliness | Slightly awkward to add LoRA on quantized weights | Smoother LoRA on quantized weights |
| Calibration data needs | 128-1024 samples | 128-512 samples |
Which One to Choose
- Chasing ultimate accuracy + compute to spare: GPTQ + Marlin kernel;
- Chasing deployment speed + LoRA fine-tuning: AWQ (the current community default);
- H100+ GPUs: FP8 (no quantization algorithm needed — native support, no accuracy loss);
- Chasing extreme compression: W4A16 + 2:4 sparsity (stacked with Pruning and Sparsification).
5. Mixed-Precision Combinations: A Selection Matrix
Different combinations of weight quantization and activation quantization address different bottleneck profiles:
| Combination | Weights | Activations | Memory Compression | Decode Speed | Prefill Speed | Accuracy Loss | Best For |
|---|---|---|---|---|---|---|---|
| W16A16 (FP16 baseline) | FP16 | FP16 | 1× | 1× | 1× | 0 | Baseline |
| W8A16 | INT8 | FP16 | 2× | ~1.5× | ~1× | <0.5% | Moderate compression, accuracy first |
| W4A16 | INT4 | FP16 | 4× | ~3-4× | ~1× | 1-3% | Decode mainstream |
| W8A8 | INT8 | INT8 | 2× | ~2× | ~2× | 1-5% | Prefill-heavy |
| W4A8 | INT4 | INT8 | 4× (weights) + 2× (activations) | ~3× | ~2× | 3-8% | Balanced but risky accuracy |
| W4A4 | INT4 | INT4 | 4×+4× | ~4× | ~4× | 5-15% | Extreme compute, poor accuracy |
| FP8 W8A8 | FP8 | FP8 | 2× | ~2× | ~2× | <1% | H100+ flagship |
| FP4 W4A4 | FP4 | FP4 | 4× | ~4× | ~4× | 2-5% | B200 Blackwell |
Decode Favors W4A16; Prefill Favors W8A8
- W4A16 is the de facto standard for online LLM serving today — decode hits the bandwidth wall, weight quantization gives a direct 4× speedup, and FP16 activations don't hurt compute (99% of compute is idle anyway);
- W8A8 suits prefill-heavy scenarios (such as offline batch processing) — prefill is compute-bound, and compressing activations to INT8 as well lets Tensor Core INT8 compute (989 TF × 2 = 1978 TOPS) run at full tilt;
- W4A4 is extreme but risky — INT4 activations are very hard to keep accurate on LLMs; use with caution.
6. Mixed-Precision Strategy: Which Layers Stay FP16
Not every layer survives INT4 quantization without accuracy loss. Sensitive layers must stay in FP16 — that is the essence of mixed precision:
| Layer | Quantization Advice | Reason |
|---|---|---|
| Final LayerNorm | FP16 | The output distribution is extremely sensitive to quantization |
| Attention softmax | FP16 | Quantization badly distorts probability distributions |
| Embedding | FP16 or INT4 | Small quantization impact; decide based on memory |
| LM head (output projection) | FP16 | Produces the output logits; quantization skews the next-token distribution |
| Q/K/V projection | INT4 OK | Intermediate representations; error is absorbed by later layers |
| FFN matmul | INT4 OK | Same as above |
Mixed Precision Is Not a "Shortcut" — It's a Necessity
Quantizing all layers uniformly to INT4 often blows up on 70B+ models (5%+ accuracy degradation). Keeping 3-5 sensitive layers in FP16 is nearly lossless (<1% loss). Both AutoGPTQ and AutoAWQ support layer-wise mixed-precision configuration — just prepare a YAML file.
7. The Marlin Kernel: Extreme Fusion for W4A16
The Pain Point
A naive W4A16 implementation:
text
1. Read INT4 weights from HBM → SRAM
2. Unpack INT4 → INT8 → FP16 (inside SRAM)
3. Run the matmul in FP16
4. Write back to HBMEvery step happens in SRAM — but written as three independent kernels for steps 1+2+3, the intermediate tensor of each step must be written back to HBM and read back in, and the dequantization benefit is eaten by HBM round trips.
What Marlin Does
Marlin (Frantar et al., 2024) fuses "read INT4 → unpack → dequantize → matmul" into a single kernel:
- Weights are dequantized on the fly inside SRAM — one HBM read, never written back;
- Warp-level tiling + shared-memory tiling — SRAM fully utilized;
- Async memcpy — HBM reads and SRAM compute run in a pipelined overlap;
- SIMD-based INT4 → FP16 conversion — one CUDA instruction unpacks 8 INT4 values into 8 FP16 values.
Result: W4A16 matmul on A100/H100 goes from ~30 TFLOPS in the naive implementation to ~150-200 TFLOPS — a 5-7× improvement.
Mainstream Implementations
- GPTQ Marlin: vLLM's default kernel for GPTQ-quantized models;
- AWQ Marlin: used by vLLM for AWQ models, 30%+ faster than AWQ's native kernel;
- exllama: an early INT4 kernel, superseded by Marlin;
- TensorRT-LLM kernels: NVIDIA's in-house implementation — performance close to Marlin but closed-source.
See Kernel Fusion and Custom Kernels.
8. Trade-offs
- GPTQ or AWQ: comparable accuracy, but AWQ kernels are faster — the mainstream pick is AWQ; choose GPTQ + Marlin for ultimate accuracy;
- 4 or 8 bits: decode favors W4A16; prefill-heavy favors W8A8; go straight to FP8 on H100+;
- Group size 32/64/128: smaller is more accurate but carries more metadata — 128 is the mainstream for INT4; INT8 doesn't need grouping;
- Mixed precision or not: strongly recommended for 70B+ models — keep the few sensitive layers (final norm / softmax / lm_head) in FP16;
- Stack LoRA or not: quantized model + LoRA is the sweet spot for fine-tuning LLMs — AWQ + LoRA has better compatibility; see Tuning and Performance Optimization.
Further Reading
- Model Quantization Fundamentals — general principles and the algorithm landscape
- Kernel Fusion and Custom Kernels — the low-level principles of the Marlin kernel
- The GPU Memory Hierarchy and the Bandwidth Wall — why weight-only quantization is a decode superpower
- Pruning and Sparsification — stacking sparse INT4 weights
- Knowledge Distillation — the training-side weapon of the compression trio
- vLLM and PagedAttention — the industrial implementation of AWQ/GPTQ Marlin
- llama.cpp and GGUF — GGUF quantization on CPU/edge
- TensorRT and GPU Inference — FP8/INT8 in the NVIDIA stack
References
- Frantar et al. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (ICLR 2023) — the original GPTQ paper
- Lin et al. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024) — the original AWQ paper
- Frantar & Alistarh. Marlin: Mixed-Precision Auto-Regressive Parallel Inference Engine (2024) — the Marlin kernel paper
- AutoGPTQ GitHub — GPTQ engineering implementation
- AutoAWQ GitHub — AWQ engineering implementation
- vLLM Quantization Documentation — using the various quantization formats in vLLM