Appearance
Classic Papers in Depth
One-sentence positioning: these ten papers, strung together, are the evolutionary history of the "modern LLM inference system" — from FlashAttention in 2022 correcting the operator bottleneck from FLOPs back to IO, to DeepSeek-V3 in 2024 pushing MoE inference serving to a hundred billion parameters, each solving a specific pain point that was then stalling an entire engine, and then changing the field.
1. Why These Ten Papers
First lay the ten on a timeline and you will see a complete history of inference acceleration:
text
2022.06 FlashAttention v1 operator-layer IO awareness, attention 2-4x faster
2022.10 GPTQ second-order quantization, 4-bit weight quantization
2022.11 SmoothQuant activation smoothing, making W8A8 viable
2022.12 Orca iteration-level scheduling, predecessor of continuous batching
2023.06 vLLM / PagedAttention KV cache managed like virtual memory, 2-4x throughput
2023.07 FlashAttention-2 parallelism and work partitioning, long-seq GPU util 35% -> 50-73%
2023.11 Speculative Decoding speculative decoding formalized, lossless
2023.12 AWQ activation-aware salient-weight protection, 3x faster than GPTQ
2024.01 Medusa multi-head parallel prediction, simplifying spec decoding
2024.06 EAGLE-2 dynamic draft tree, 70%+ acceptance rate
2024.12 DeepSeek-V3 671B MoE + MTP, open-source benchmark| Paper | First author | Venue | One-line contribution |
|---|---|---|---|
| FlashAttention | Dao | NeurIPS 2022 | IO-aware attention operator: tiling + online softmax, 2-4x faster on A100 |
| GPTQ | Frantar | ICLR 2023 | Post-training 4-bit weight quantization with Hessian information |
| SmoothQuant | Xiao | ICML 2023 | Smooth activation outliers into the weights — the icebreaker of W8A8 |
| Orca | Yu | OSDI 2022 | Iteration-level scheduling, the predecessor of continuous batching |
| vLLM / PagedAttention | Kwon | SOSP 2023 | KV cache managed like virtual memory, block table + block-level allocation |
| FlashAttention-2 | Dao | 2023 | Rebuilt parallelism and work partitioning, 50-73% long-seq GPU utilization |
| Speculative Decoding | Leviathan | 2023 | Small model guesses + large model verifies in parallel, lossless |
| AWQ | Lin | MLSys 2024 | Activation-aware salient-weight protection, 3x faster than GPTQ |
| Medusa | Cai | ICML 2024 | Main-model multi-head parallel prediction, simplifying spec decoding |
| EAGLE-2 | Li | EMNLP 2024 | Dynamic draft tree, acceptance rate pushed to 70%+ |
| DeepSeek-V3 | DeepSeek-AI | 2024 | 671B MoE + MTP, the open-source MoE inference benchmark |
Why these ten and not others? Three reasons:
- Each represents a paradigm switch: FlashAttention corrected the bottleneck from FLOPs to IO; vLLM redefined LLM serving from "loading a model" to "managing memory"; GPTQ/AWQ upgraded quantization from "uniform compression" to "differentiated protection"; EAGLE upgraded speculative decoding from "an independent small model" to "drafting from the main model's hidden state." Read these ten switches and you grasp the decade's main intellectual thread.
- Each is sized for a close read: except for the DeepSeek-V3 Technical Report, the bodies are mostly 10-15 pages, with method and experiments focused on a single core idea — far more readable than today's giant technical reports.
- Each has been repeatedly validated by industry: FlashAttention entered PyTorch main; vLLM is the de facto inference engine; GPTQ/AWQ are HuggingFace quantize defaults; EAGLE is integrated in SGLang/vLLM. Read them and you have read the tools you use every day.
Before reading this page
Read Start Here first for the section's three routes, then build a timeline with the Paper Map. Every close-reading below follows a fixed four-step structure — problem background -> core mechanism -> experimental conclusions -> implications for today — so you can take notes as you go.
2. FlashAttention v1: The Operator-Layer IO-Awareness Milestone (Dao, NeurIPS 2022)
2.1 Problem Background
Attention in Transformers is O(n²) complexity, but the mainstream approach from 2017-2022 was "reduce FLOPs" — sparse attention (Longformer, Big Bird), linear attention (Performer), and so on. Tri Dao at Stanford asked the opposite question: is attention slow really because of too many FLOPs? He profiled and found that standard attention on an A100 has arithmetic intensity of only about 10 FLOPs/byte, far below the roofline knee (~100 FLOPs/byte) — the bottleneck is HBM bandwidth, not FLOPs. Specifically: standard attention must read and write three O(n²) intermediate matrices in HBM — S = QKᵀ, P = softmax(S), O = PV — each hundreds of megabytes, saturating the bandwidth over and over.
2.2 Core Mechanism: Tiling + Online Softmax + No Write-Back of Intermediates
FlashAttention's core idea condenses into one sentence: move attention's intermediate state from HBM into SRAM. But SRAM is only about 192 KB per SM — too small for an entire n×n matrix — so tiling is mandatory. The difficulty: softmax needs to see the whole row to compute the denominator; how do you guarantee correctness with only tiles?
text
Standard attention:
Q, K, V in R^{n x d} live in HBM
Loop 1: compute S = QK^T (write back to HBM, an n x n matrix)
Loop 2: P = softmax(S) (read S, write P back to HBM)
Loop 3: O = PV (read P, write O back to HBM)
-> HBM read/write volume O(n^2 + n·d) times several
FlashAttention:
Q, K, V in R^{n x d} live in HBM
Outer loop over K/V blocks (block size B_r x B_c, sized to SRAM):
load the Q block, K block, and V block into SRAM
compute S_block = Q_block · K_block^T in SRAM
use online softmax to update the numerator and denominator incrementally:
m_new = max(m_old, rowmax(S_block))
f_new = exp(m_old - m_new) * f_old + sum(exp(S_block - m_new))
compute O_block += exp(S_block - m_new) * V_block in SRAM
finally normalize O and write it back to HBM
-> HBM read/write volume O(n^2·d/M), where M is the SRAM sizeThree key mechanisms:
- Blocked loops: the outer loop is over K/V blocks (not Q blocks — that is what v2 improves), and the inner loop computes attention between the Q block and the K/V blocks.
- Online softmax (from Milakov & Gimelshein 2018): computing softmax block by block does not require seeing the whole row — you only maintain "the current maximum m" and "the weighted sum f," updating incrementally as each new block arrives. This is the key mathematical innovation.
- No write-back of intermediates: S and P never leave SRAM; only the final result O is written back. No n×n intermediate matrix lands in HBM during the forward pass. To save memory during the backward pass, instead of checkpointing, it recomputes — running the forward once more to compute gradients, which is faster than storing the intermediate state.
2.3 Experimental Conclusions
- Training GPT-2 (sequence 1024): 2.7x faster than PyTorch's standard attention; at sequence 4096, 7.8x.
- Wall-clock on GPT-2 training: HBM read/write volume drops from O(n²) to O(n²·d/M), measured IO reduced 2-4x.
- Memory: drops from O(n²) to O(n), long-sequence training no longer OOMs.
- Accuracy: exact attention, no approximation whatsoever.
2.4 Implications for Today
- Profile before optimizing: FlashAttention was "disruptive" because it first profiled "the bottleneck is IO" and then prescribed the cure. Measure first, then optimize — the iron law of systems optimization.
- Recomputation can be cheaper: not storing the intermediate state and recomputing it in the backward pass is faster than storing it — this defies the "avoid recomputation" common sense but matches the real bandwidth ratio between IO and SRAM. Today's optimization paradigm has been rewritten by this idea. See Kernel Fusion and Custom Kernels and GPU Architecture and Optimization.
- Mechanisms generalize: FlashAttention's "tiling + online reduction" idea was generalized beyond attention (Flash-Decoding, FlashInfer's block sparsity, fused FFN). A model of one trick changing an entire field.
A reminder about one number
"2-4x" is a "right here, right now" number: change the batch, the sequence length, or the A100 hardware and it no longer holds. What is truly worth taking away is the mechanistic judgment — "the bottleneck is IO; tiling + online softmax + recomputation" — not that speedup number.
3. FlashAttention-2: From Usable to Good (Dao, 2023)
3.1 Problem Background
FlashAttention v1 turned attention from "cannot compute" into "can compute," but Dao himself quickly found that v1's GPU utilization on long sequences was only about 35% — still not feeding the H100/A100 enough. The bottleneck was v1's loop order: the outer loop over K/V blocks caused Q blocks to be reloaded repeatedly; meanwhile v1 partitioned work by "thread block" without considering warp-level load balance, so warp utilization was uneven on long sequences.
3.2 Core Mechanism: Swapping the Loop Order + Warp-Level Work Partitioning
- Swap the loop order: v2 makes the outer loop over Q blocks and the inner loop over K/V blocks. Now a Q block is loaded only once, and the K/V blocks iterate in SRAM — the "output" O for that Q block accumulates in SRAM and is not written back to HBM until the Q block is done.
- Warp-level work partitioning: v1 had one thread block compute one row of attention; v2 splits a thread block into 4 warps, each computing a different K/V block, and merges them with a warp-level reduction — making better use of the GPU's hardware parallelism.
- Reduce non-matmul FLOPs: v2 optimized the softmax and rescale non-matmul compute, raising the matmul share from ~70% in v1 to 80%+, closer to the Tensor Core sweet spot.
3.3 Experimental Conclusions
- On A100: 1.5x faster than v1 at sequence 1024, 2x faster at sequence 4096.
- GPU utilization: from 35% in v1 to 50-73% (of theoretical peak).
- On H100: through later porting, FA2 became the de facto standard for attention on H100.
3.4 Implications for Today
- The same paper can iterate on itself: v1 found the mechanism, v2 found the engineering headroom — mechanism first, engineering follow-up is the standard rhythm of research.
- The GPU programming model must reach warp level: the three-level thread block / warp / thread programming model recurs throughout systems papers; anyone doing inference optimization must know it cold.
- Continuous profiling drives continuous optimization: after v1 Dao still profiled, found the 35% utilization, and produced v2. No profiling, no optimization.
The "relay" between v1 and v2
v1 found the core mechanism of "tiling + online softmax"; v2 solved "how to feed that mechanism to the GPU." Read v1 for intuition, v2 for engineering detail. v3 then added FP8 + asynchrony on H100; the three form a complete evolution chain. Read alongside FlashInfer to see "how one idea evolves across hardware generations."
4. GPTQ: Second-Order Quantization (Frantar, ICLR 2023)
4.1 Problem Background
In 2022 LLM quantization research, naive round-to-nearest collapsed in accuracy at 4-bit, and LLM.int8() only reached 8-bit. Could a 175B model keep its accuracy at 4-bit? Frantar et al. (IST Austria) found an idea in the OBQ (Optimal Brain Quantizer) from the 1990s: minimize the quantization error by quantizing "column by column," updating the remaining columns to compensate after each column is quantized. But OBQ was infeasible for a 175B model — computing a Hessian had to be made affordable.
4.2 Core Mechanism: Layer-Wise + Cholesky + Lazy Updates
GPTQ makes three engineering modifications on top of OBQ:
text
1. Layer-wise (not whole-model)
Quantize only each layer's W -> Q; each layer is optimized independently.
One layer = one independent quantization problem.
2. Ordered quantization + error compensation
For each column j:
q_j = argmin_q || W_j - q ||_{H^{-1}} (weighted by Hessian information)
W_remaining <- W_remaining - W_j * q_j (compensate the error into the remaining columns)
H = X X^T is the layer's Hessian (outer product of input activations)
3. Cholesky makes the Hessian tractable
H can be ill-conditioned on a 175B model; GPTQ stabilizes the numerics with a Cholesky decomposition.
"Lazy updates" batch the inter-column updates to reduce GPU synchronization.
4. Estimate H with a small calibration set
128 samples suffice to estimate H; no training data needed.The core idea in one sentence: use Hessian information to quantize columns, and spread each column's error across the remaining columns. This upgrades "which weight columns matter" from "looking at weight magnitudes" to "looking at second-order activation information."
4.3 Experimental Conclusions
- OPT-175B at 4-bit: accuracy loss <1% (C4 perplexity 0.32 higher than FP16).
- Quantization speed: about 4 hours for a 175B model on 4 GPUs (prior schemes were infeasible).
- Memory: a 175B model drops from 350 GB (FP16) to about 100 GB (4-bit), runnable on a single A100 80 GB (with offload).
4.4 Implications for Today
- Second-order information is the gold standard of quantization: every W4 quantization scheme today (AWQ, SpinQuant) compares against GPTQ — GPTQ is the paradigm's starting point for second-order quantization.
- Old statistical methods still work on new problems: OBQ was a 1990s pruning method, and GPTQ "rediscovered" and adapted it for LLMs — standing on another field's shoulders is the shortest path to lowering the cost of innovation.
- The "engineering nature" of quantization papers: GPTQ's core algorithm is not hard; the hardness is engineering — Cholesky, lazy updates, batch processing. For inference-optimization papers, engineering details often decide production readiness more than the algorithm itself. Quantization mechanics: Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
GPTQ's engineering limits
GPTQ's Hessian estimation needs calibration data and is sensitive to the choice of it; accuracy varies with batch size; and the GPU implementation needs dedicated kernels (such as AutoGPTQ). Algorithm + engineering + kernel — all three in one, none dispensable.
5. SmoothQuant: The Icebreaker of Activation Quantization (Xiao, ICML 2023)
5.1 Problem Background
GPTQ solved 4-bit weights, but activation quantization (W8A8) kept failing on LLMs: LLM activations have extreme outliers (about 0.1% of channels have activations 100x those of other channels), INT8 quantization overflows on these outliers, and W8A8 collapses on models >6.7B. How can activations be quantized too? SmoothQuant's answer: don't quantize the activations directly — migrate their difficulty to the weights.
5.2 Core Mechanism: Per-Channel Scaling = Difficulty Migration
text
Original: Y = X · W X has outliers, W is normal
INT8(X) overflows
Smooth: Y = (X · s) · (W / s) (s is a per-channel scaling factor)
= X' · W' X' = X·s flattens the outliers
W' = W/s enlarges the corresponding weight channels
now INT8(X') does not overflow, and INT8(W') accuracy is still acceptableKey insight: matrix multiplication is invariant to per-channel scaling — multiplying a column of X by s and dividing the corresponding row of W by s leaves Y unchanged. But the distribution of X changes (smoother) and the distribution of W changes (some channels larger), and both become more suitable for INT8 quantization.
The choice of s: $s_j = \max(|X_j|)^\alpha / \max(|W_j|)^{1-\alpha}$, with α usually 0.5, splitting the difficulty fifty-fifty between weights and activations.
5.3 Experimental Conclusions
- OPT-175B and BLOOM-176B at W8A8: accuracy nearly identical to FP16 (perplexity gap < 0.2).
- Speed: W8A8 is 1.5-2x faster than FP16 on A100 (using INT8 Tensor Cores).
- Integration: SmoothQuant has entered PyTorch, HuggingFace, vLLM, and other mainstream frameworks.
5.4 Implications for Today
- Difficulty migration is a general idea: "borrow" the difficulty of the hard-to-quantitize side (activations) to the easy side (weights), at low cost and high gain. This "equivalent transform + difficulty redistribution" was inherited by later rotation quantization (QuaRot, SpinQuant).
- Activation quantization is far harder than weight quantization: weight distributions are relatively stable, while activations swing violently with the input; outliers are the core problem of LLM quantization. After SmoothQuant, every activation-quantization scheme must answer "how do you handle outliers."
- The triumph of the "equivalent transform" idea: from SmoothQuant (scaling) to QuaRot (rotation), the core breakthroughs of quantization lie not in "smarter rounding" but in "smarter equivalent transforms" — a paradigm shift in quantization research.
SmoothQuant and GPTQ are complementary
GPTQ solves weight quantization to 4-bit; SmoothQuant solves activation quantization to 8-bit. A mainstream inference-engine combination today: weights 4-bit (GPTQ/AWQ) + activations 8-bit (SmoothQuant) = W4A8. They do not conflict and are often deployed together.
6. Orca: The Predecessor of Continuous Batching (Yu, OSDI 2022)
6.1 Problem Background
In early 2022, the mainstream way to serve LLMs was "request-level batching" — batch several requests together and return only when the whole batch finishes. But LLM generation is variable-length: one request may generate 8 tokens, another 800. Under request-level batching, short requests must wait for long requests to finish before being released, and the GPU idles heavily. Orca at OSDI 2022 asked: can each request in a batch be scheduled independently as soon as it generates one token?
6.2 Core Mechanism: Iteration-Level Scheduling
text
Traditional request-level batching:
T0: [Req A 8 steps] [Req B 50 steps] [Req C 100 steps] --> batch
T1-T100: the whole batch runs 100 steps; all requests finish together
<- short requests are dragged down by long requests
Orca iteration-level scheduling:
T0: [Req A 8 steps] [Req B 50 steps] [Req C 100 steps] --> batch
T8: Req A finishes, releases resources; new request Req D joins the batch
T9-T50: the batch keeps running; requests join and leave mid-way
<- the GPU almost never idles; throughput = sum(request rate)/average stepsKey designs:
- Iteration-level scheduling: the unit of scheduling is "one forward pass" (generating one token), not "one complete request." A request can be released or join after completing any token.
- Selective batching: different requests make their forward pass in the same iteration, but each request's KV cache is managed independently.
- Scheduler: FIFO + admission control decides who enters and who leaves the next iteration.
6.3 Experimental Conclusions
- Versus request-level batching: throughput improves 2-37x (depending on the tail of the load).
- Tail latency: significantly reduced under mixed long/short requests (short requests are no longer held back by long ones).
- Compute utilization: GPU utilization rises from ~30% to 70%+.
6.4 Implications for Today
- The scheduling unit decides resource efficiency: switching the unit from "request" to "iteration" rewrote the throughput model of LLM serving. Every mainstream engine today (vLLM, SGLang, TensorRT-LLM) uses iteration-level batching — Orca defined the scheduling paradigm of LLM serving.
- The value of "system-level innovation" in OSDI/SOSP papers: Orca proposed no new algorithm but a new scheduling unit. The core of a systems paper is "redefining the boundary of the problem" — its biggest difference from an algorithms paper.
- Serving mechanics: see Batching and Request Scheduling and Model Serving and Orchestration.
The relationship between Orca and vLLM
Orca's innovation was the "scheduling unit"; vLLM's innovation was "memory management." The two are complementary — today's vLLM implementation uses both Orca's iteration-level scheduling and PagedAttention. Read Orca to understand the "scheduling" half; read vLLM to understand the "memory" half.
7. vLLM / PagedAttention: The LLM Serving Revolution (Kwon, SOSP 2023)
7.1 Problem Background
Orca solved the "scheduling unit," but LLM serving had another killer: KV cache memory fragmentation. In early 2023 the vLLM team at Berkeley profiled and found that when serving LLaMA-13B, the KV cache occupied 30+ GB of memory, and because requests differ in length, about 60-80% of memory was wasted on internal and external fragmentation. Could the operating-system idea of "virtual memory + paging" be applied to the KV cache?
7.2 Core Mechanism: Manage the KV Cache Like Virtual Memory
text
Traditional KV cache:
each request pre-allocates a contiguous max_length of memory
-> internal fragmentation (actual < max): 80%+ wasted
-> external fragmentation (gaps between requests): 30%+ wasted
PagedAttention:
the KV cache is split into fixed-size "blocks" (e.g. 16 tokens)
each request's logical KV is mapped to physical blocks via a block table
+-------------+
| Logical KV | --+
| block 0 | | block table (like an OS page table)
| block 1 | +--▶ [physical block 3] [physical block 7] [physical block 12]...
| block 2 | |
+-------------+ |
allocate new blocks on demand; no pre-allocation of max_length
Attention computation:
the kernel must do "logically contiguous" attention, but the blocks are physically non-contiguous
-> the kernel indexes by the block table internally, like an OS page-table translationThree key mechanisms:
- Block-level KV cache: split the KV cache into fixed-size blocks (block size usually 16), allocate on demand, and eliminate internal fragmentation (wasting at most block_size-1 tokens).
- Block table: each request maintains a logical-block-to-physical-block mapping table, like an OS page table.
- PagedAttention kernel: the attention operator must access KV in logical order while the blocks are physically non-contiguous — the kernel indexes by the block table internally. This is the hardest part of the engineering.
- Copy-on-write: in scenarios like beam search, multiple candidates can share the KV cache and copy a new block only on write — saving memory.
7.3 Experimental Conclusions
- Throughput: 2-4x over HuggingFace Transformers; 1.5-2x over Orca-style serving.
- Memory utilization: from ~20-40% to 80%+ (fragmentation eliminated).
- Tail latency: P99 latency reduced by 50%+ (especially under high load).
- Beam search: copy-on-write lets multiple candidates share KV, saving 50%+ of memory.
7.4 Implications for Today
- Borrowing ideas across disciplines: vLLM transplanted the operating system's 50-year-old "virtual memory + paging" idea into GPU memory — old ideas can still open new ground on new problems. A model of how systems papers borrow ideas.
- Kernel engineering decides feasibility: the core difficulty of PagedAttention is the kernel — a non-contiguous KV attention operator indexed by the block table, which the vLLM team spent months writing. Algorithm + engineering + kernel: all three in one to ship.
- Open source decides impact: beyond the paper, the team's open-source repo became the de facto standard — paper + open source + community in one let vLLM surpass Orca and FasterTransformer to become the de facto LLM serving engine. Case: vLLM and PagedAttention.
A frequently asked vLLM interview question
"Is PagedAttention mathematically equivalent to standard attention?" Answer: yes. Only the physical layout differs (contiguous vs. blocked); mathematically it computes the same softmax(QKᵀ)V — the kernel just indexes by the block table internally. An architecture choice is essentially a physical-layout choice.
8. Speculative Decoding: The Formalization of Speculative Decoding (Leviathan, 2023)
8.1 Problem Background
The core bottleneck of autoregressive LLM generation: generating one token requires recomputing the entire model's forward pass. Leviathan et al. (Google/Technion) asked: can a small model first guess k tokens, and the large model verify all k in one parallel forward pass? If the guesses are right, it is as if k tokens were generated in one step; if wrong, at least the correct part is kept. Theoretically lossless (the output distribution matches the large model's).
8.2 Core Mechanism: Draft + Verify + Accept/Reject
text
Draft phase:
a small model q (1B) autoregressively generates k tokens: x_1, x_2, ..., x_k
Verify phase:
the large model p (70B) computes the true distributions p(x_i | x_<i) at x_1, ..., x_k in one forward pass
Accept/Reject (based on rejection sampling):
for each x_i:
if q(x_i | x_<i) <= p(x_i | x_<i): accept, move to the next
otherwise: accept with probability (p-q)/p; else reject and resample from p
Theoretical guarantee:
the final output distribution = p's distribution, lossless
Where the speedup comes from:
the small model is fast + the large model verifies k in parallel = ~2-3x faster than large-model autoregression
(depends on the acceptance rate)Key insights:
- The acceptance rate decides the speedup: the more accurate the small model, the higher the acceptance rate and the bigger the speedup. Typical acceptance rates are 50-70%.
- Theoretically lossless: rejection sampling guarantees the final distribution matches the large model's exactly — no accuracy loss of any kind.
- Draft model choice: it can be a smaller version of the same model (e.g. Llama-70B paired with Llama-7B) or the early layers of the same model shared.
8.3 Experimental Conclusions
- On Chinchilla-70B with a Chinchilla-7B draft: 2-3x speedup (text generation).
- Acceptance rate: 50-70% (task-dependent; higher on code tasks).
- Accuracy: exactly consistent with large-model autoregression (KL divergence 0).
8.4 Implications for Today
- "Small cost to guess + large cost to verify" is a general idea: speculative decoding is not unique to LLMs — CPU branch prediction, cache prefetching, and normalized edit distance are all variants of the same idea.
- Theoretical losslessness is an engineering victory: accelerating 2-3x without losing accuracy is the benchmark of a systems paper. Lossless vs. lossy is a key watershed of systems papers.
- The direction of follow-up work: Medusa replaced "an independent small model" with "extra heads on the main model"; EAGLE replaced "token-level drafting" with "hidden-state drafting"; DeepSeek-V3 MTP learned "drafting ability at training time." Speculative decoding is one of the most active research directions of LLM inference in 2024-2025. See Speculative Decoding and Medusa/EAGLE.
The "cost ledger" of speculative decoding
Speculative decoding is not a free lunch: drafting costs extra compute, and verification still computes a k-step forward pass. Gain = acceptance rate x (large-model step cost / small-model step cost) - verification overhead. When the acceptance rate < 30%, speculative decoding can be slower than plain autoregression — which is why the EAGLE series works so hard to raise the acceptance rate. If an interview asks "when does speculative decoding beat plain autoregression," answer: if and only if acceptance rate x k > verification overhead / step cost.
9. AWQ: Activation-Aware Salient-Weight Protection (Lin, MLSys 2024)
9.1 Problem Background
GPTQ protects important weights with Hessian information — effective but slow (4 hours to quantize a 175B model). Lin et al. at MIT/Hugging Face asked: is there a faster way to reach GPTQ's accuracy? They observed a counter-intuitive phenomenon: only 0.1-1% of the weights are critical to model accuracy, and these "salient weights" can be located via activation statistics — no Hessian needed.
9.2 Core Mechanism: Activation-Aware + Salient Weights + Scale Protection
text
Observation 1: only ~0.1-1% of a large model's weights are "salient"
(critical to model accuracy)
Observation 2: salient weights != weights with large magnitudes
they are the weights corresponding to channels with large activations
-> locate salient weights with activation statistics (per-channel max of X)
Method:
1. Use 128 samples to compute per-channel activation maxima and find the top-1% "salient channels"
2. Scale-protect the salient channels' weights:
W' = W * s, where s > 1 for salient channels
during quantization, s "amplifies" these weights so the INT4 rounding error is relatively smaller
3. On dequantization, W = W'/s restores the original
4. Equivalent to: salient weights get higher quantization precision than ordinary weightsKey insights:
- Salient weights != large weights: GPTQ locates them indirectly with a Hessian; AWQ locates them directly with activations — activations are the most direct signal of "which weights matter."
- Scale protection: through the equivalent transform (W' = W*s, X' = X/s) the quantization error of salient weights becomes relatively smaller — the reverse application of the SmoothQuant idea.
- No backpropagation: AWQ needs no gradients, only a small amount of activation statistics, so it quantizes fast.
9.3 Experimental Conclusions
- Llama-7B/13B/70B at 4-bit: accuracy comparable to GPTQ (perplexity gap < 0.1).
- Quantization speed: 3x faster than GPTQ (no Hessian, no gradients).
- Inference speed: with INT4 kernels, 2-3x faster than FP16.
- Integration: the default 4-bit scheme of HuggingFace transformers, vLLM, and TensorRT-LLM.
9.4 Implications for Today
- Statistics can replace second-order information: AWQ replaces the Hessian with activation maxima — an order of magnitude cheaper in engineering cost at comparable accuracy — a victory of engineering optimization.
- Sparse protection is a general paradigm: from GPTQ's Hessian to AWQ's activations, the idea is "find the few important parameters and protect them." Every later quantization scheme inherits this idea.
- The "engineering-friendly" dimension of quantization papers: GPTQ's algorithm is elegant but engineering-heavy; AWQ's algorithm is simple and engineering-friendly — in the competition for production adoption, simple and fast often win.
Common misconceptions about GPTQ vs AWQ
GPTQ and AWQ are comparable at 4-bit (gap < 0.1 perplexity), but they protect important weights differently: GPTQ uses the Hessian indirectly + error compensation; AWQ uses activations directly + scale protection. A common interview question is "when to use GPTQ and when to use AWQ": answer — AWQ suits fast quantization and default deployment; GPTQ suits pursuing extreme accuracy (5-bit) or offline scenarios.
10. Medusa: Multi-Head Parallel Prediction (Cai, ICML 2024)
10.1 Problem Background
Speculative Decoding uses an independent small model for drafting, but training a small model is inconvenient (training it again) and the small model's features do not align perfectly with the large model's. Cai at Princeton/DeepSeek asked: can the main model grow its own "multiple heads" to draft, without an independent small model? That is Medusa.
10.2 Core Mechanism: Main Model + Multiple Decoding Heads
text
Standard LLM:
main model hidden state h_t -> LM head (vocab V) -> predicts x_{t+1}
Medusa:
main model hidden state h_t -> LM head -> predicts x_{t+1}
-> Medusa head 1 -> predicts x_{t+2}
-> Medusa head 2 -> predicts x_{t+3}
...
-> Medusa head k -> predicts x_{t+k+1}
a Medusa head is a simple MLP taking h_t and outputting a distribution over V
at training time the main model is frozen; only the Medusa heads are trained (a few hours)
Verify phase:
one forward pass computes all heads' predictions
tree-based attention verifies the k candidates at once
accept the highest-probability contiguous sequenceKey insights:
- No independent small model: Medusa heads train quickly (a few hours), have few parameters (each head is an MLP), and share the main model's hidden state — saving all the trouble of a small model.
- Tree-based attention: verifying multiple candidate token trees in one forward pass is Medusa's contribution to engineering speculative decoding.
- Trading a little losslessness for speed: Medusa does not strictly guarantee losslessness (acceptance is heuristic rather than rejection sampling), but the real-world accuracy loss is almost unmeasurable and the speedup is significant.
10.3 Experimental Conclusions
- Vicuna-13B + Medusa-2: 2.3-2.8x speedup (over plain autoregression).
- Training cost: about 1-2 hours per head (on ~1B-scale data).
- Accuracy: almost imperceptible loss on evaluations such as MT-Bench.
10.4 Implications for Today
- "Main model + lightweight attachment" is an engineering-friendly paradigm: compared with training an independent small model, letting the main model grow "extra heads" is extremely engineering-friendly — Medusa's biggest contribution to the engineering of speculative decoding.
- The trade-off between lossless and engineering-friendly: Medusa gives up strict losslessness but gains engineering simplicity and training speed. In engineering optimization, "lossless" is not the only goal; "good enough and simple" often wins.
- Medusa is the prelude to EAGLE: Medusa's "main-model multi-head" idea was inherited and upgraded by EAGLE — EAGLE replaced "multiple heads" with "hidden-state drafting," a step further. Only by understanding Medusa can you understand EAGLE's evolution.
Why Medusa's acceptance rate is not high
Medusa's typical acceptance rate is about 30-40%, lower than the EAGLE series. The reason: a Medusa head sees only h_t (a single hidden state) yet must predict t+2, t+3, ... — not enough information. EAGLE makes the draft model a "lightweight transformer layer" that can see multiple hidden states, so the acceptance rate is much higher. The evolution more information -> multiple hidden states -> higher acceptance rate is the main line of speculative decoding.
11. EAGLE-2: Dynamic Draft Trees and the Acceptance-Rate Breakthrough (Li, EMNLP 2024)
11.1 Problem Background
Medusa uses a static tree shape (k fixed heads) for drafting, but an LLM's "predictability of the next token" is context-dependent — some tokens are easy to guess ("the", "of"), some are hard (the first character of a person's name). Can the draft tree deform dynamically with the context? EAGLE-2 (Peking/Tencent) answers: yes.
11.2 Core Mechanism: Hidden-State Drafting + Dynamic Trees + Context-Aware Pruning
text
EAGLE-1 baseline:
draft model = one transformer layer, taking the main model's h_t and x_t, outputting x_{t+1}
compared with Medusa it adds "feature alignment with h_t"
acceptance rate 50%+
EAGLE-2 upgrade:
static tree -> dynamic tree:
compute the cumulative probability of every path in real time during drafting
sort by cumulative probability and take the top-k paths
cumulative probability = prod p(x_i | x_<i) under the draft model's distribution
key: the tree's shape adapts to the context
easy-to-predict context -> a deeper, narrower tree (one main path going deep)
hard-to-predict context -> a wider, shallower tree (many branches, high acceptance)
+ tree-shaped attention verifies in one forward passKey insights:
- Static tree vs. dynamic tree: Medusa has k fixed heads; EAGLE-2 prunes dynamically by probability — handing the draft tree's shape to the context.
- Cumulative probability as the pruning criterion: prune low-probability paths early and give the verification budget to high-probability paths — an application of tree-search ideas to speculative decoding.
- Hidden-state alignment: the draft model takes h_t, not just x_t, as input; the feature space aligns with the main model, and the acceptance rate rises significantly.
11.3 Experimental Conclusions
- Acceptance rate: 70%+ (EAGLE-1 was 50%, Medusa 30-40%).
- Speedup: 3x+ on MT-Bench, 5x+ on some tasks.
- Training cost: the draft model trains in 1-2 days (ready out of the box after integration into vLLM/SGLang).
11.4 Implications for Today
- Tree search + probability pruning is a general idea: EAGLE-2 applies the "probability pruning" of MCTS to speculative decoding — old idea + new problem = new paper.
- The acceptance rate decides everything: every improvement in speculative decoding ultimately attacks "the acceptance rate" number. Medusa 30% -> EAGLE-1 50% -> EAGLE-2 70% -> EAGLE-3 80%+ — every 10-point rise in acceptance rate lifts the speedup by a tier.
- Integration decides adoption: EAGLE is integrated in both vLLM and SGLang — paper + engine integration = impact, an important difference between systems papers and algorithms papers. See Speculative Decoding and Medusa/EAGLE.
The evolution logic of the EAGLE series
EAGLE-1 found the "hidden-state drafting" mechanism; EAGLE-2 found the "dynamic tree" engineering optimization; EAGLE-3 found the "deeper features + training objective" algorithmic improvement. Three generations evolve in the order "mechanism -> engineering -> algorithm," exactly the same as the FlashAttention three generations (mechanism -> engineering -> hardware co-design) — the general law of how systems papers iterate on themselves.
12. DeepSeek-V3 Technical Report: The New Benchmark of MoE Inference Serving (DeepSeek-AI, 2024)
12.1 Problem Background
The open-source LLMs of early 2024 were Llama-3-70B (dense) and Mixtral-8x22B (8-expert MoE). DeepSeek asked: can MoE push open-source models to a hundred billion parameters while costing less to serve than Llama-3-70B? DeepSeek-V3 answers: a 671B-total / 37B-active MoE + MTP (multi-token prediction).
12.2 Core Mechanism: Auxiliary-Loss-Free MoE + MTP Training
text
Model architecture:
61 Transformer layers
shared expert + 256 routed experts
each token activates 8 experts (1 shared + 7 routed)
671B total parameters, 37B active
128K context
Training:
14.8T tokens (high-quality synthetic + real data)
FP8 training (hardware: 2048 H800s)
auxiliary-loss-free MoE: uses a bias term for load balancing without losing accuracy
MTP: the model predicts multiple tokens at once at training time, learning draft ability
Inference:
MTP used as speculative decoding: the main model has built-in draft ability, no independent small model needed
FP8 inference + PagedAttentionKey innovations:
- Auxiliary-loss-free MoE: traditional MoE forces load balancing with an auxiliary loss, which costs accuracy. DeepSeek-V3 uses a bias term for "soft" load balancing, lossless to accuracy.
- MTP training objective: predict k tokens at once at training time, training speculative-decoding ability into the main model so it is used directly as the draft at inference.
- FP8 training + inference: FP8 training on H800s, and FP8 for inference too — the benchmark case of FP8 engineering for open-source large models.
- Training cost: about 2 months on 2048 H800s, totaling about $5.6 million — an order of magnitude cheaper than closed-source models like GPT-4.
12.3 Experimental Conclusions
- Performance: near Llama-3-405B on benchmarks such as MMLU, GSM8K, and MATH, surpassing it on some tasks.
- Inference cost: $0.27 per million input tokens, 4x lower than Llama-3-70B (by performance-per-cost).
- Price war: after DeepSeek-V3 launched, OpenAI and Anthropic lowered API prices by about 30-50% in the same period.
12.4 Implications for Today
- MoE is the future of open-source large models: DeepSeek-V3 proved that "open source + MoE + high-quality data" can compete with closed-source dense models at lower inference cost. Since 2025, almost all mainstream open-source large models are MoE.
- The training-time internalization of speculative decoding: DeepSeek-V3 treats speculative decoding as a training objective (MTP), not an inference-time add-on — this is the paradigm shift of speculative-decoding research, and 2025 papers are all following.
- The technical report is the new form of systems paper: DeepSeek-V3 is not an academic-conference paper but a technical report (arXiv 2412.19437), yet the industry treats it as "required reading" — the venues of systems papers are diversifying: technical report + open weights = impact.
- The rise of Chinese open-source power: DeepSeek-V3 marks a Chinese team reaching world-class in large-model systems optimization, with a major impact on the global open-source ecosystem and commercial pricing.
Open questions of DeepSeek-V3
DeepSeek-V3's MoE serving poses new deployment challenges: memory residency of 256 experts, load balancing of expert routing, and cross-node communication — all far more complex than Llama-3-70B. It took SGLang, vLLM, and other engines half a year to get efficient DeepSeek-V3 inference working. New model architectures vs. inference-engine capability is always a catch-up race.
13. The Right Way to Read a Paper: The Three-Pass Method
A close read is not "reading word by word from start to finish once" but reading three times with rhythm; each pass has a different goal and output. The method comes from S. Keshav's classic paper How to Read a Paper:
| Pass | Time budget | What to read | Output |
|---|---|---|---|
| Pass 1 - bird's-eye | 5-10 min | Title, abstract, introduction, conclusion, captions of all figures | Can answer: what problem is solved? The method in one sentence? The effect? |
| Pass 2 - structure | ~1-2 hours | Full body, figures, pseudocode outline, method flow | Can draw the method flowchart and retell the key experiments and ablations |
| Pass 3 - close read | Hours to days | Line-by-line derivation, kernel reimplementation, critical review | Can point out limitations, propose improvements, and judge "how it would fare in my setting" |
For inference-systems papers, the three-pass method has three extra points:
- Always read abstract -> profile figures -> conclusion first. A systems paper's "profile figures" are extremely dense: one roofline figure is worth three paragraphs, and one batch-throughput curve exposes the method's true level. Look at the figures first, then read the text with questions in hand.
- In Pass 2, always do "ablation tracking." Wherever the paper shows "Table X: Ablation," stop and ask: what did the authors prove by removing this component? This is the core method for distinguishing "genuine innovation" from "engineering stacking." FlashAttention doubling its numbers without an IO-awareness ablation is far less credible than "remove IO-awareness and the speed drops to 1.1x."
- In Pass 3, translate every "speedup" into a "boundary." Attach a qualifying condition to every 2x and 3x: under what model, what batch, what sequence length, and what hardware does it hold? Once translated, you have truly read the paper (more methodology in Reading Discipline & FAQ).
14. The Common Laws of the Ten Papers
Spanning 2022-2024, across the five areas of operators, quantization, serving, decoding, and MoE, the ten papers are strikingly consistent when placed together:
Law 1: Good papers come from solving a concrete bottleneck, not chasing a new concept.
- FlashAttention solved "the attention bottleneck is IO, not FLOPs"; GPTQ solved "4-bit quantization accuracy collapse"; SmoothQuant solved "activation outliers make W8A8 infeasible"; Orca solved "request-level batching wastes the GPU"; vLLM solved "KV cache fragmentation"; Speculative Decoding solved "autoregression produces one token per step"; AWQ solved "GPTQ is too slow"; Medusa solved "an independent small model is hard to train"; EAGLE-2 solved "low acceptance rate"; DeepSeek-V3 solved "high inference cost of open-source dense models."
- Every pain point was a concrete obstacle blocking the road at the time; in hindsight all are "so important they need no explanation." Pain points define problems, and problems define innovation.
Law 2: The core mechanism is simple enough to state in one sentence.
- FlashAttention: tiling + online softmax + no write-back of intermediates; GPTQ: quantize columns with Hessian information + error compensation; SmoothQuant: per-channel scaling migrates difficulty from activation to weight; vLLM: paged KV cache management; Speculative Decoding: small model guesses + large model verifies in parallel; AWQ: locate salient weights with activations + scale protection; Medusa: main-model multi-head parallel prediction; EAGLE-2: dynamic draft tree + probability pruning; DeepSeek-V3: auxiliary-loss-free MoE + MTP. Mechanisms that truly change a field can usually be stated in one sentence.
Law 3: All use experiments to prove "it is the mechanism at work," not just to report results.
- FlashAttention uses HBM read/write volume to prove "the bottleneck is IO"; GPTQ contrasts Hessian vs. uniform quantization to prove "second-order information matters"; vLLM compares memory utilization before and after fragmentation to prove "paging is necessary"; EAGLE-2 ablates static vs. dynamic trees to prove "dynamics contribute acceptance rate." Good papers do not just report results — they design experiments that "prove the results are caused by the mechanism."
Law 4: The breakthroughs almost all come from "borrowing ideas across fields."
- FlashAttention borrows the IO-complexity analysis of databases; vLLM borrows the operating system's virtual memory; GPTQ borrows 1990s pruning; SmoothQuant borrows equivalent transforms from statistics; Speculative Decoding borrows CPU branch prediction and rejection sampling; EAGLE-2 borrows MCTS probability pruning; DeepSeek-V3 borrows speculative decoding but internalizes it into training. Standing on another discipline's shoulders is the shortest path to lowering the cost of innovation.
Law 5: Each opens a door rather than closing one.
- FlashAttention made "IO-aware operators" a new paradigm; vLLM made "memory management" a serving standard; GPTQ/AWQ made "4-bit quantization" the default; EAGLE made "speculative decoding" mainstream; DeepSeek-V3 made "open-source MoE + training-time MTP" a new direction. A simple test of a paper's value: does it make follow-up work easier, cheaper, and more likely? Apply this test to reread the ten papers and each passes solidly.
Summing up in one sentence
Reading inference papers is less about acquiring knowledge than about learning how to find and solve systems bottlenecks. The shared methodology of the ten is: find a concrete, profileable bottleneck -> give a simple, explainable solution -> use ablations to prove the causal chain between the mechanism and the speedup. If you can reproduce this methodology, you have the skeleton to do systems optimization independently. The rest is walking every coordinate on the Paper Map.
Further Reading
- Start Here — the entry to the papers section and the conceptual foundation; set your reading goal first
- Reading Paths — specific paper lists and reading orders for the three routes
- Paper Map — place these ten papers in the decade-wide macro coordinate system
- Frontier Advances — after the classics, walk toward the 2024-2026 frontier
- What Is Inference Acceleration? — the conceptual foundation of inference deployment
- A Brief History — place the papers on a larger timeline
- Case studies: vLLM and PagedAttention, TensorRT-LLM, Speculative Decoding and Medusa/EAGLE, Distributed Inference (TP/PP)
- Glossary — consult while reading
References
The following are all real public resources; the originals are directly accessible:
- Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (NeurIPS 2022)
- Dao. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning (2023)
- Frantar et al. GPTQ: Accurate Post-Training Quantization for GPT (ICLR 2023)
- Xiao et al. SmoothQuant: Accurate and Efficient Post-Training Quantization for LLMs (ICML 2023)
- Yu et al. Orca: A Distributed Serving System for Transformer-Based LLMs (OSDI 2022)
- Kwon et al. Efficient Memory Management for LLMs Serving with PagedAttention (SOSP 2023)
- Leviathan et al. Fast Inference from Transformers via Speculative Decoding (2023)
- Lin et al. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024)
- Cai et al. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (ICML 2024)
- Li et al. EAGLE-2: Faster Speculative Decoding with Dynamic Draft Trees (EMNLP 2024)
- DeepSeek-AI. DeepSeek-V3 Technical Report (2024)
- Keshav. How to Read a Paper (2007) — the original source of the three-pass method
- Tri Dao's homepage — the FlashAttention author's blog, often with author's-view follow-up notes
- vLLM project — the industrial-strength implementation of PagedAttention; issues and PRs hide a wealth of engineering detail
- FlashAttention official repo — the official implementation of the FlashAttention family and related operators