Skip to content

Weight-Only Quantization and Mixed Precision

At a glance Weight-only quantization compresses LLM weights to INT4/INT8 while keeping activations in FP16/BF16 — purpose-built for the decode-phase bandwidth wall. This article covers GPTQ's Hessian-based second-order compensation, AWQ's salient-weight protection, the Marlin kernel's extreme fusion, and how to choose among W8A16/W4A16/W4A8 combinations.

Weight-Only Quantization and Mixed Precision ​

Concept Definition: The Strongest Weapon Against the Decode-Phase Bandwidth Wall ​

LLM inference is memory-bound in the decode phase — every generated token requires reading the entire model weights from HBM once. Weight-only quantization is purpose-built for this: compress the weights to INT4/INT8 (4×/2× fewer bytes) while keeping activations in FP16/BF16 (no accuracy loss), and decode throughput gains 3-4×.

Two key insights for understanding weight-only quantization:

  1. It compresses only weights, not activations — because activations are the small share (a forward pass computes only a few KB for the current token), while weights are the big share (140GB for a 70B model);
  2. It targets decode; prefill benefits little — prefill is compute-bound, so cutting bytes doesn't directly translate into speed; decode is memory-bound, so cutting bytes does.

Weight-only quantization is the de facto standard of LLM inference today — vLLM, TensorRT-LLM, and llama.cpp all ship built-in support.

1. The Basic Recipe of Weight-Only Quantization ​

text
Before:  W (FP16) × X (FP16)  →  Y (FP16)         # decode-phase bandwidth bottleneck
Quant:   W (INT4) → dequant → W' (FP16) × X (FP16) → Y
                                       ↑
                            dequantization done on the fly before the matmul

The key question: where does dequantization happen? Two mainstream approaches:

ApproachFlowAdvantagesDisadvantages
Offline dequantizationW INT4 → W FP16 (one-shot dequant before inference)Simple to implementNo memory saved, no speedup — an anti-pattern
Online dequantization + fused kernelDequantize each weight group inside the matmul4× memory savings, 4× bandwidth savingsKernels must be hand-optimized (Marlin)

Don't "Dequantize Offline"

"Dequantize INT4 weights back to FP16 before inference" is the most common beginner trap — it saves no memory (you still store FP16) and is actually slower (an extra dequantization step). True weight-only quantization requires a fused kernel — dequantization happens inside the matmul: each INT4 group is dequantized and consumed immediately, and the weights stay in INT4 the whole time. The Marlin kernel exists precisely for this.

2. GPTQ: Hessian-Based Second-Order Compensation ​

Core Idea ​

GPTQ (Frantar et al., 2023) is one of the most accurate algorithms for INT4 weight quantization. Its core: quantize weights column by column, and after each column use Hessian information to adjust the remaining columns and compensate for the error.

The math:

text
Quantization error minimization: minimize || W × X - Q × X ||²
where X is the calibration data, and X × X^T ≈ H (the Hessian matrix)

GPTQ uses a greedy algorithm based on second-order information (the Hessian): quantize column by column, and after each column project the error onto the remaining columns weighted by the Hessian, letting subsequent columns compensate for the error already incurred. This "compensating" quantization minimizes the accumulated error.

Engineering ​

python
# Simplified AutoGPTQ workflow
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig

quant_config = BaseQuantizeConfig(
    bits=4,
    group_size=128,
    desc_act=False,           # whether to order columns by the Hessian
)
model = AutoGPTQForCausalLM.from_pretrained("llama-2-70b")
model.quantize(calibration_data, quant_config)
model.save_quantized("llama-2-70b-gptq-4bit")

Strengths and Limitations ​

  • Strengths: excellent INT4 accuracy (<1% loss), broad compatibility (almost every LLM has a GPTQ build);
  • Limitations: slow quantization (calibration + Hessian computation takes hours for a 70B model), and kernel speed slightly trails AWQ.

3. AWQ: Protecting Salient Weights ​

Core Idea ​

The core observation of AWQ (Lin et al., 2023): about 1% of the weights in an LLM are "salient" and decisive for accuracy — typically the weights on channels with large activation magnitudes. If this 1% is protected, the accuracy loss from quantizing the other 99% is negligible.

But how should "protecting 1% of weights" be implemented? AWQ uses a clever trick: equivalent scaling.

text
Original:   Y = W × X
Equivalent: Y = (W / s) × (s · X) = W' × X'
                    ↑              ↑
             scaled weights    scaled activations

By multiplying channels with large activation magnitudes by s > 1 (making the salient weights relatively smaller, landing them in an INT4-friendlier region), the corresponding channels of X are equivalently divided by s — mathematically equivalent, but the salient weights lose less during quantization.

Why AWQ Is Faster Than GPTQ ​

AWQ's quantization process is a lightweight search that needs no Hessian — it only observes activation magnitudes over the weights, and a few hundred forward passes determine the scale for every group. The kernel implementation is also more efficient — the AWQ kernel (vLLM implementation) typically leads the GPTQ Marlin by 10-30% at comparable accuracy.

Engineering ​

python
# Simplified AutoAWQ workflow
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer

model = AutoAWQForCausalLM.from_pretrained("llama-2-70b")
tokenizer = AutoTokenizer.from_pretrained("llama-2-70b")
quant_config = {"w_bit": 4, "q_group_size": 128, "zero_point": True}
model.quantize(tokenizer, quant_config)
model.save_quantized("llama-2-70b-awq-4bit")

4. GPTQ vs. AWQ: A Practical Comparison ​

DimensionGPTQAWQ
Core algorithmHessian-based second-order compensationSalient-weight protection + equivalent scaling
Quantization timeSlow (hours for 70B)Fast (30-60 minutes)
INT4 accuracy<1% loss<1% loss (comparable)
Kernel speedNear AWQ after Marlin optimizationFaster out of the box
EcosystemAutoGPTQ, vLLM, TensorRT-LLMAutoAWQ, vLLM, TensorRT-LLM
Fine-tuning friendlinessSlightly awkward to add LoRA on quantized weightsSmoother LoRA on quantized weights
Calibration data needs128-1024 samples128-512 samples

Which One to Choose

  • Chasing ultimate accuracy + compute to spare: GPTQ + Marlin kernel;
  • Chasing deployment speed + LoRA fine-tuning: AWQ (the current community default);
  • H100+ GPUs: FP8 (no quantization algorithm needed — native support, no accuracy loss);
  • Chasing extreme compression: W4A16 + 2:4 sparsity (stacked with Pruning and Sparsification).

5. Mixed-Precision Combinations: A Selection Matrix ​

Different combinations of weight quantization and activation quantization address different bottleneck profiles:

CombinationWeightsActivationsMemory CompressionDecode SpeedPrefill SpeedAccuracy LossBest For
W16A16 (FP16 baseline)FP16FP161×1×1×0Baseline
W8A16INT8FP162×~1.5×~1×<0.5%Moderate compression, accuracy first
W4A16INT4FP164×~3-4×~1×1-3%Decode mainstream
W8A8INT8INT82×~2×~2×1-5%Prefill-heavy
W4A8INT4INT84× (weights) + 2× (activations)~3×~2×3-8%Balanced but risky accuracy
W4A4INT4INT44×+4×~4×~4×5-15%Extreme compute, poor accuracy
FP8 W8A8FP8FP82×~2×~2×<1%H100+ flagship
FP4 W4A4FP4FP44×~4×~4×2-5%B200 Blackwell

Decode Favors W4A16; Prefill Favors W8A8

  • W4A16 is the de facto standard for online LLM serving today — decode hits the bandwidth wall, weight quantization gives a direct 4× speedup, and FP16 activations don't hurt compute (99% of compute is idle anyway);
  • W8A8 suits prefill-heavy scenarios (such as offline batch processing) — prefill is compute-bound, and compressing activations to INT8 as well lets Tensor Core INT8 compute (989 TF × 2 = 1978 TOPS) run at full tilt;
  • W4A4 is extreme but risky — INT4 activations are very hard to keep accurate on LLMs; use with caution.

6. Mixed-Precision Strategy: Which Layers Stay FP16 ​

Not every layer survives INT4 quantization without accuracy loss. Sensitive layers must stay in FP16 — that is the essence of mixed precision:

LayerQuantization AdviceReason
Final LayerNormFP16The output distribution is extremely sensitive to quantization
Attention softmaxFP16Quantization badly distorts probability distributions
EmbeddingFP16 or INT4Small quantization impact; decide based on memory
LM head (output projection)FP16Produces the output logits; quantization skews the next-token distribution
Q/K/V projectionINT4 OKIntermediate representations; error is absorbed by later layers
FFN matmulINT4 OKSame as above

Mixed Precision Is Not a "Shortcut" — It's a Necessity

Quantizing all layers uniformly to INT4 often blows up on 70B+ models (5%+ accuracy degradation). Keeping 3-5 sensitive layers in FP16 is nearly lossless (<1% loss). Both AutoGPTQ and AutoAWQ support layer-wise mixed-precision configuration — just prepare a YAML file.

7. The Marlin Kernel: Extreme Fusion for W4A16 ​

The Pain Point ​

A naive W4A16 implementation:

text
1. Read INT4 weights from HBM → SRAM
2. Unpack INT4 → INT8 → FP16 (inside SRAM)
3. Run the matmul in FP16
4. Write back to HBM

Every step happens in SRAM — but written as three independent kernels for steps 1+2+3, the intermediate tensor of each step must be written back to HBM and read back in, and the dequantization benefit is eaten by HBM round trips.

What Marlin Does ​

Marlin (Frantar et al., 2024) fuses "read INT4 → unpack → dequantize → matmul" into a single kernel:

  • Weights are dequantized on the fly inside SRAM — one HBM read, never written back;
  • Warp-level tiling + shared-memory tiling — SRAM fully utilized;
  • Async memcpy — HBM reads and SRAM compute run in a pipelined overlap;
  • SIMD-based INT4 → FP16 conversion — one CUDA instruction unpacks 8 INT4 values into 8 FP16 values.

Result: W4A16 matmul on A100/H100 goes from ~30 TFLOPS in the naive implementation to ~150-200 TFLOPS — a 5-7× improvement.

Mainstream Implementations ​

  • GPTQ Marlin: vLLM's default kernel for GPTQ-quantized models;
  • AWQ Marlin: used by vLLM for AWQ models, 30%+ faster than AWQ's native kernel;
  • exllama: an early INT4 kernel, superseded by Marlin;
  • TensorRT-LLM kernels: NVIDIA's in-house implementation — performance close to Marlin but closed-source.

See Kernel Fusion and Custom Kernels.

8. Trade-offs ​

  • GPTQ or AWQ: comparable accuracy, but AWQ kernels are faster — the mainstream pick is AWQ; choose GPTQ + Marlin for ultimate accuracy;
  • 4 or 8 bits: decode favors W4A16; prefill-heavy favors W8A8; go straight to FP8 on H100+;
  • Group size 32/64/128: smaller is more accurate but carries more metadata — 128 is the mainstream for INT4; INT8 doesn't need grouping;
  • Mixed precision or not: strongly recommended for 70B+ models — keep the few sensitive layers (final norm / softmax / lm_head) in FP16;
  • Stack LoRA or not: quantized model + LoRA is the sweet spot for fine-tuning LLMs — AWQ + LoRA has better compatibility; see Tuning and Performance Optimization.

Further Reading ​

References ​