Appearance
Paper Map
One-sentence positioning: this map lays the key inference-acceleration and deployment papers of 2014-2025 along the two axes of theme and year, forming a coordinate system you can search at any time — from PipeDream's pipeline parallelism in 2014 to EAGLE-3's 6.5x speedup in 2025. After reading it, you can state "what bottleneck each key paper solved, which papers it connects to before and after, and whether it deserves a close read or a skim."
The biggest obstacle to reading papers is not "failing to understand one" but "not knowing which one to read next." Tutorials teach in order, but the world of papers is organized by problems: the answer to the same problem ("how to keep attention from being stuck on HBM") was FlashAttention v1 in 2022, FlashAttention-2 in 2023, FlashAttention-3 in 2024 (H100 FP8), and also vAttention in 2024 (using software PagedAttention to approximate hardware-native sparse attention). If you read only along the timeline, you get lost in "who optimized whom"; if you read only by theme, you miss the macro narrative of "why hardware-generation shifts reshape algorithmic routes." The purpose of this map is to overlay theme and time so that, facing any paper, you can answer three questions: which technical line does it belong to? where does it sit in the paradigm evolution? and which papers form a "must-read chain" with it?
How the map divides work with the other pages
This papers section has five pages with different roles: Start Here explains "why read inference papers," Reading Paths explains "how to read by goal," this page (Paper Map) explains "what papers in the field are worth reading and how they relate," Classic Papers in Depth digs into papers one by one, and Frontier Advances tracks the new trends of 2024-2026. In one sentence: the map sets coordinates, close-reading gives depth, and the frontier points the direction.
1. How to Read This Map
Before unfolding the map, three conventions:
- Numbering convention: every paper is given an arXiv number (or the formal venue), so the full text is free at https://arxiv.org/abs/NUMBER; older papers without an arXiv entry are marked "—" with the venue given, and the original links are in the references.
- Year convention: a paper's year always means its first public release (formal venue publication, or arXiv preprint submission). For example, the FlashAttention v1 preprint was submitted in June 2022 and it appeared at NeurIPS 2022, so the table records 2022.
- Deep-read convention: the map gives only a "one-line contribution" — an answer to "why is this paper worth remembering." To master a paper, jump to Classic Papers in Depth; to chase its follow-ups, use Frontier Advances and the search method in FAQ.
Use the three coordinates together: theme decides which line it belongs to, year decides where it sits in the paradigm evolution, and importance decides how much time you invest. Here is the overview of the map:
text
Theme dimension (six technical lines)
┌────────────────────────────────────┐
│ 1. Operator optimization │
│ 2. Quantization & compression │
Time dimension | │ 3. Serving systems & scheduling │
│ 4. Speculative & parallel decoding │
│ 5. Distributed training/inference │
│ 6. On-device & edge deployment │
└────────────────────────────────────┘2. The Academic Map by Theme
2.1 Operator Optimization: From Naive Attention to IO-Awareness and Hardware Co-Design (2017-2025)
The core question of this group is always: why is attention's O(n²) complexity the bottleneck — is it a FLOPs bottleneck or an IO bottleneck? After the Transformer in 2017, the community first worked on FLOPs (sparse attention, Longformer, Big Bird) until FlashAttention pointed out in 2022 that the bottleneck was HBM bandwidth, correcting the route back to "IO-awareness + operator fusion."
| Paper | Authors/Institution | Year | Venue/arXiv | One-line contribution |
|---|---|---|---|---|
| Attention Is All You Need | Vaswani et al. (Google Brain) | 2017 | arXiv:1706.03762 | Proposes the Transformer and the pure-attention architecture — the source of O(n²) complexity |
| Longformer / Big Bird | Beltagy et al. / Zaheer et al. | 2020 | arXiv:2004.05150 / arXiv:2007.14062 | Sparse attention replacing global attention with "local window + a few global tokens," O(n) compute on long documents |
| FlashAttention v1 | Dao et al. (Stanford) | 2022 | arXiv:2205.14135 | IO-aware attention operator: tiling + online softmax + no write-back of intermediates to HBM, 2-4x faster on A100 |
| FlashAttention-2 | Dao (Stanford) | 2023 | arXiv:2307.08691 | Rebuilt parallelism and work partitioning, raising GPU utilization on long sequences from 35% to 50-73% |
| FlashAttention-3 | Shah et al. (Princeton) | 2024 | arXiv:2407.08608 | FP8 + asynchrony (warp-specialized) on H100, 1.2 PFLOPs peak, 1.5-2x faster than FA2 |
| FlashInfer | Ye et al. (Princeton/DeepSeek) | 2024 | arXiv:2401.12448 | Block-sparse + composable KV formats, unifying paged KV and variable-length attention in one kernel |
| vAttention | Prabhu et al. (Georgia Tech) | 2024 | arXiv:2405.04447 | Achieves software PagedAttention's function with hardware-native sparse attention, avoiding SM stalls |
Why read the FlashAttention series before vAttention
The three FlashAttention generations (v1→v2→v3) show the evolution paradigm of "the same math formula, three generations of engineering": v1 found the bottleneck (IO), v2 optimized parallelism, v3 co-designed with hardware (H100 FP8). vAttention asks "can hardware sparse attention replace software PagedAttention," a new-paradigm debate from 2024. Read this line and you grasp the core law of "how operator optimization follows hardware generations." Operator mechanics in detail: Kernel Fusion and Custom Kernels.
2.2 Quantization and Compression: From Uniform Quantization to Second-Order Information and Rotation (2020-2025)
The core question of quantization: can weights and activations be represented with fewer bits at acceptable accuracy loss? This line went through three paradigm shifts in the 2020s: naive round-to-nearest → LLM.int8() discovers outliers → GPTQ/AWQ/SmoothQuant protect important parameters more cleverly → rotation quantization (QuaRot/SpinQuant) and FP8/FP4 hardware quantization in 2024.
| Paper | Authors/Institution | Year | Venue/arXiv | One-line contribution |
|---|---|---|---|---|
| LLM.int8() | Dettmers et al. (Washington) | 2022 | arXiv:2208.07339 | Discovers "outlier features" in models >6.7B and proposes mixed precision (outlier columns in FP16, the rest INT8) — the start of LLM quantization research |
| GPTQ | Frantar et al. (IST Austria) | 2022 | arXiv:2210.17323 | Post-training 4-bit weight quantization with Hessian information, <1% accuracy loss on 175B models — the second-order paradigm |
| SmoothQuant | Xiao et al. (MIT/ByteDance) | 2022 | arXiv:2211.10438 | Smooths activation outliers into the weights, making W8A8 quantization possible — the icebreaker of activation quantization |
| AWQ | Lin et al. (MIT/Hugging Face) | 2023 | arXiv:2306.00978 | Activation-aware salient-weight protection, 3x faster than GPTQ with comparable accuracy — a common production choice |
| ZeroQuant | Yao et al. (Microsoft) | 2022 | arXiv:2206.01861 | Joint weight-and-activation quantization + per-block calibration, an early industrial post-training scheme |
| FP8 LLM Inference (NVIDIA) | NVIDIA | 2023 | NVIDIA FP8 white paper | H100 introduces FP8 Tensor Cores — a hardware-level sweet-spot format |
| QuaRot | Ashkboos et al. (ETH) | 2024 | arXiv:2404.00456 | Rotation quantization: random orthogonal transforms "flatten" outliers, greatly improving W4A4 accuracy |
| SpinQuant | Liu et al. (Meta) | 2024 | arXiv:2405.16406 | Learned rotation matrices, better than QuaRot's random rotation, W4A4 approaching W4A16 |
Two threads of quantization
Weight side: GPTQ and AWQ use second-order/activation information to locate "salient weights" — "differentiated protection." Activation side: SmoothQuant migrates difficulty from activation to weight, and QuaRot/SpinQuant "flatten" the distribution with rotation. The two threads converge in 2024 at W4A4 + rotation, letting 4-bit inference approach 16-bit accuracy. Quantization mechanics in detail: Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
2.3 Serving Systems and Scheduling: From Naive Batching to Prefill/Decode Disaggregation (2022-2025)
The core question of serving systems: with many users and variable-length requests, how do you keep the GPU from idling, memory from fragmenting, and tail latency under control? Orca proposed iteration-level scheduling in 2022, vLLM managed the KV cache like virtual memory in 2023, DistServe/Splitwise physically disaggregated prefill and decode in 2024, and SARATHI balanced compute with chunked prefill.
| Paper | Authors/Institution | Year | Venue/arXiv | One-line contribution |
|---|---|---|---|---|
| Orca | Yu et al. (Microsoft) | 2022 | OSDI 2022 | Iteration-level scheduling — requests of different lengths advance in the same batch, the predecessor of continuous batching |
| vLLM / PagedAttention | Kwon et al. (Berkeley) | 2023 | SOSP 2023 | Manage the KV cache like virtual memory with a block table and block-level allocation, 2-4x throughput |
| AlpaServe | Li et al. (Berkeley) | 2023 | OSDI 2023 | Model parallelism + time multiplexing, using statistical multiplexing to guide multi-model deployment |
| SARATHI | Agarwal et al. (Microsoft) | 2023 | arXiv:2308.16369 | Chunked prefill: split a long prefill into chunks that share a batch with decode, balancing compute and memory |
| DistServe | Zhong et al. (Peking/Tsinghua) | 2024 | arXiv:2401.09670 | Physically separate prefill and decode onto different GPUs, each scaling independently |
| Splitwise | Patel et al. (Microsoft/IIT) | 2024 | HPCA 2024 | Similar to DistServe but focused on cluster-level "phase-specific machine" scheduling |
| SGLang | Zheng et al. (Berkeley/DeepSeek) | 2024 | arXiv:2312.07104 | RadixAttention: a prefix tree reuses the KV cache — a throughput weapon for multi-turn dialogue and few-shot scenarios |
| Mooncake | Moonshot AI | 2024 | arXiv:2407.00079 | A planet-scale serving architecture that disaggregates the KV cache, treating it as a first-class citizen |
Three stages of serving-system evolution
Stage 1 (Orca, 2022): swap "batching by request" for "batching by iteration" — a throughput leap. Stage 2 (vLLM, 2023): swap "pre-allocated memory" for "on-demand block-level allocation" — KV fragmentation eliminated. Stage 3 (DistServe/SARATHI, 2024): disaggregate "prefill + decode on the same GPU in the same batch," because their compute/memory ratios are exactly opposite. Read these three stages and you understand the history of LLM serving. Serving mechanics: Batching and Request Scheduling and Model Serving and Orchestration.
2.4 Speculative and Parallel Decoding: From Single-Step to Multi-Token Prediction (2022-2025)
The core bottleneck of the decode layer: autoregressive generation produces one token per step, recomputing the KV cache 1000+ times just to generate 1000 tokens. Speculative decoding's idea is "guess k tokens at small cost, then verify them with the large model in one pass." Leviathan formalized it in 2023, Medusa/EAGLE-1/2 each upgraded drafting in 2024, and EAGLE-3 pushed the speedup to 6.5x in 2025.
| Paper | Authors/Institution | Year | Venue/arXiv | One-line contribution |
|---|---|---|---|---|
| Speculative Decoding | Leviathan et al. (Google/Technion) | 2023 | arXiv:2211.17192 | A small model guesses + the large model verifies in parallel, lossless acceleration — the formalization of speculative decoding |
| SpecInfer | Miao et al. (CMU/SCU) | 2023 | arXiv:2302.02018 | Token-tree parallel verification + lion-hearted SSG, upgrading speculative decoding with multiple drafts |
| Medusa | Cai et al. (Princeton/DeepSeek) | 2024 | arXiv:2401.10774 | The main model predicts in parallel with multiple heads, no independent small model needed — simplifying speculative decoding |
| EAGLE-1 | Li et al. (Peking/Tencent) | 2024 | arXiv:2401.15077 | Drafts from the main model's hidden state, aligning draft and verify features, 50%+ acceptance rate |
| EAGLE-2 | Li et al. (Peking/Tencent) | 2024 | arXiv:2406.16858 | Dynamic draft tree that prunes the tree by context probability, pushing the acceptance rate to 70%+ |
| EAGLE-3 | Li et al. (Peking/Tencent) | 2025 | arXiv:2503.01840 | Deeper hidden-layer features + an improved training objective, 3-6.5x speedup |
| SpecForge | SGLang team | 2025 | SGLang blog | A multi-draft-model training pipeline, treating speculative decoding as an "engineering product" rather than a "model product" |
| DeepSeek-V3 MTP | DeepSeek-AI | 2024 | arXiv:2412.19437 | The main model natively supports multi-token prediction, learning speculative-decoding ability at training time |
The "cost ledger" of speculative decoding
The core trade-off of speculative decoding is: the extra compute of drafting vs. the steps saved by verification. The more accurate the draft and the higher the acceptance rate, the more is saved; the larger the draft, the higher the overhead. The evolution logic of the EAGLE series is "use the main model's hidden state to make the draft both accurate and small," while DeepSeek-V3's MTP takes another route — teach the main model to emit multiple tokens at once at training time, internalizing speculative decoding into the training objective. Which one becomes mainstream is an open question for 2025-2026. Cases: Speculative Decoding and Medusa/EAGLE.
2.5 Distributed Training/Inference: From Data Parallelism to Expert and Pipeline Parallelism (2019-2025)
Distributed is no longer "a training-time thing" — MoE inference makes expert parallelism a hard requirement for deployment, and DeepSeek-V3's 671B total / 37B active parameters make inference-side distributed scheduling as complex as training-side.
| Paper | Authors/Institution | Year | Venue/arXiv | One-line contribution |
|---|---|---|---|---|
| Megatron-LM | Shoeybi et al. (NVIDIA) | 2019 | arXiv:1904.10509 | Tensor parallelism: shard each layer's matrix multiplications by column/row across GPUs — infrastructure for training large models |
| PipeDream | Narayanan et al. (Stanford/Google) | 2019 | arXiv:1806.03312 | Pipeline parallelism + micro-batching + 1F1B scheduling, making "layer-sharding across GPUs" viable |
| MegaScale | Jiang et al. (ByteDance) | 2024 | NSDI 2024 | An industrial system for trillion-parameter MoE training, with communication/compute overlap and fault recovery |
| Mixtral of Experts | Mistral AI | 2024 | arXiv:2401.04088 | Open-weights 8-expert MoE, bringing MoE from closed to open source — expert parallelism becomes standard on the inference side |
| DeepSeek-V3 | DeepSeek-AI | 2024 | arXiv:2412.19437 | 256 experts + 8 active MoE + MTP, pushing open-source MoE inference serving to a hundred-billion-parameter benchmark |
The training/inference boundary of distributed is blurring
Megatron and PipeDream are training-side papers, but their parallel schemes (TP/PP/SP) were inherited directly by inference — multi-GPU inference in TensorRT-LLM and vLLM uses the same sharding. MoE inference (Mixtral, DeepSeek-V3) brought "expert parallelism" from training into inference, forcing inference engines to add new scheduling. Today you cannot do inference deployment without understanding training-side parallelism. See Distributed Inference (TP/PP).
2.6 On-Device and Edge Deployment: Cramming an LLM into Phones and Laptops (2020-2025)
As LLMs move onto phones and laptops, the question shifts from "how to be faster" to "how to run at all in 4-8 GB of memory." This line includes on-device architectures (MobileBERT), on-device inference engines (MLC-LLM, llama.cpp), and on-device quantization (MCT-LLM).
| Paper | Authors/Institution | Year | Venue/arXiv | One-line contribution |
|---|---|---|---|---|
| MobileBERT | Sun et al. (HUST/Google) | 2020 | arXiv:2004.02584 | Distillation + bottleneck structures layer by layer, compressing BERT to run on a phone — the start of on-device Transformers |
| MCT-LLM | Linux Foundation / ONNX Format | 2023 | Microsoft Build 2023 | Microsoft's on-device INT4 quantization, running Llama-7B on a Surface Laptop |
| MLC-LLM | TVM team (OctoML/Apache) | 2023 | mlc.ai | Compile LLMs to GPU/NPU/CPU with TVM — a compiler approach to on-device LLMs |
| llama.cpp | Gerganov et al. | 2023 | github.com/ggerganov/llama.cpp | Pure C++ + the GGUF quantization format, running Llama on CPUs and low-end GPUs — the on-device de facto standard |
Two routes for on-device inference
The compiler route (MLC-LLM): treat an LLM as a computation graph and compile it to different hardware backends with TVM — elegant engineering but slow to adapt. The direct-implementation route (llama.cpp): rewrite the whole inference path in pure C++ + SIMD — flexible but adaptation is manual. Today llama.cpp is the on-device de facto standard thanks to its simplicity and active community, but whether the compiler route overtakes it in the NPU era is undecided. See Mobile Deployment and llama.cpp and GGUF.
3. Paradigm-Shift Timeline: A Decade of Evolution in One Diagram
Arrange the key papers of the six themes along a timeline and you can see the rhythm of paradigm shifts in inference acceleration:
text
2017 Transformer -- the source of O(n^2) complexity
|
2019 Megatron-LM/PipeDream -- parallel infrastructure for training large models
|
2020 Longformer/BigBird -- naive sparse attention (sidestepping O(n^2))
|
2022 LLM.int8() -- outliers discovered; quantization research breaks through
| GPTQ -- second-order quantization
| SmoothQuant -- activation smoothing
| Orca -- iteration-level scheduling
| FlashAttention v1 -- IO-aware operator (paradigm turning point)
|
2023 vLLM/PagedAttention -- KV cache managed like virtual memory (paradigm turning point)
| AWQ -- activation-aware quantization
| FlashAttention-2 -- parallelism and work-partition optimization
| Speculative Decoding -- speculative decoding formalized
| llama.cpp/MLC-LLM -- on-device LLM lands
| SARATHI -- chunked prefill
|
2024 FlashAttention-3 -- H100 FP8 co-design
| vAttention -- hardware-native sparsity replaces software paged
| DistServe/Splitwise -- prefill/decode physically disaggregated
| SGLang -- RadixAttention prefix-tree reuse
| Medusa/EAGLE-1/2 -- speculative decoding upgrades
| QuaRot/SpinQuant -- rotation quantization
| Mixtral/DeepSeek-V3 -- MoE inference serving at scale
| Mooncake -- KV-cache-disaggregated architecture
|
2025 EAGLE-3 -- 6.5x speculative decoding
| FP4/NVFP4 -- Blackwell B200 era
| P-EAGLE/SpecForge -- speculative decoding productionized
| ...The two paradigm turning points are the most worth remembering: FlashAttention v1 in 2022 corrected the bottleneck from FLOPs back to IO, and vLLM in 2023 redefined LLM serving from "loading a model" to "managing memory." Prioritize close-reading papers from these two years, because they defined today's engineering coordinate system.
4. Task-Based Paper Index
Different roles and scenarios care about different papers. This table lets you look up papers by "task":
| What you are doing | Must-read papers | Advanced papers |
|---|---|---|
| Writing an attention kernel | FlashAttention v1/v2 | FlashAttention-3, FlashInfer, vAttention |
| Optimizing LLM serving throughput | vLLM, Orca | SARATHI, DistServe, Splitwise, SGLang, Mooncake |
| Quantizing a 70B model for launch | GPTQ, AWQ | SmoothQuant, LLM.int8(), QuaRot, SpinQuant |
| Speeding up the decode phase | Speculative Decoding, EAGLE-2 | Medusa, EAGLE-1/3, SpecInfer |
| Deploying an MoE model | Mixtral, DeepSeek-V3 | MegaScale, Megatron-LM |
| Running on-device LLMs | llama.cpp, MLC-LLM | MobileBERT, MCT-LLM |
| Squeezing performance on H100/B200 | FlashAttention-3, FP8 inference | NVFP4, QuaRot, vAttention |
| Cramming before an interview | All of Classic Papers in Depth | Pair with Frontier Advances |
5. Reading by Importance Tier
If time is short, pick papers by the tiers below:
Three-tier reading list
Must-read 5 (anyone doing inference deployment should read these)
- FlashAttention v1 (2022)
- vLLM / PagedAttention (2023)
- GPTQ or AWQ (pick one)
- EAGLE-2 (2024)
- Orca or SARATHI (pick one)
Strongly recommended 5 (systematic deep dive or job hunting)
- FlashAttention-2 (2023)
- SmoothQuant (2022)
- Medusa (2024)
- DistServe (2024)
- FlashAttention-3 (2024)
Researcher must-read 5 (frontier tracking / paper writing)
- vAttention (2024)
- SGLang (2024)
- QuaRot or SpinQuant (2024)
- DeepSeek-V3 Technical Report (2024)
- EAGLE-3 (2025)
6. Further Reading
- Start Here — the entry to the papers section and the three values of reading papers
- Reading Paths — arrange the papers in this map into three executable routes
- Classic Papers in Depth — section-by-section breakdowns of the must-read 5 plus selected strongly-recommended papers
- Frontier Advances — a systematic sort of the important 2024-2026 breakthroughs
- Reading Discipline & FAQ — methodology for reading papers
- What Is Inference Acceleration? — the conceptual foundation
- A Brief History — place the papers on a larger timeline
- Glossary — consult while reading
- Case studies: vLLM and PagedAttention, TensorRT-LLM, Speculative Decoding and Medusa/EAGLE, Distributed Inference (TP/PP), Mobile Deployment
References
The following are all real public resources; the originals are directly accessible:
- Vaswani et al. Attention Is All You Need (NeurIPS 2017)
- Beltagy et al. Longformer: The Long-Document Transformer (2020)
- Zaheer et al. Big Bird: Transformers for Longer Sequences (NeurIPS 2020)
- Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (NeurIPS 2022)
- Dao. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning (2023)
- Shah et al. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision (2024)
- Ye et al. FlashInfer: Efficient and Customizable Attention Engine for LLMs (2024)
- Prabhu et al. vAttention: Efficient Attention with Suffix Parallelism (2024)
- Dettmers et al. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (NeurIPS 2022)
- Frantar et al. GPTQ: Accurate Post-Training Quantization for GPT (ICLR 2023)
- Xiao et al. SmoothQuant: Accurate and Efficient Post-Training Quantization for LLMs (ICML 2023)
- Lin et al. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024)
- Ashkboos et al. QuaRot: Outlier-free 4-bit Inference for Rotated LLMs (2024)
- Liu et al. SpinQuant: LLM Quantization with Learned Rotations (2024)
- Yu et al. Orca: A Distributed Serving System for Transformer-Based LLMs (OSDI 2022)
- Kwon et al. Efficient Memory Management for LLMs Serving with PagedAttention (SOSP 2023)
- Li et al. AlpaServe: Scaling for Large-Scale Deep Learning Serving (OSDI 2023)
- Agarwal et al. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills (2023)
- Zhong et al. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving (2024)
- Patel et al. Splitwise: Efficient generative LLM inference using phase splitting (HPCA 2024)
- Zheng et al. SGLang: Efficient Execution of Structured Language Model Programs (2024)
- Mooncake. Mooncake: A KV-Centric Decoupled Architecture for LLM Serving (2024)
- Leviathan et al. Fast Inference from Transformers via Speculative Decoding (2023)
- Miao et al. SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference (SOSP 2023)
- Cai et al. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (ICML 2024)
- Li et al. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (2024)
- Li et al. EAGLE-2: Faster Speculative Decoding with Dynamic Draft Trees (EMNLP 2024)
- Li et al. EAGLE-3: Scaling up Inference Speed (2025)
- DeepSeek-AI. DeepSeek-V3 Technical Report (2024)
- Shoeybi et al. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism (2019)
- Narayanan et al. PipeDream: Generalized Pipeline Parallelism for DNN Training (SOSP 2019)
- Jiang et al. MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs (NSDI 2024)
- Jiang et al. Mixtral of Experts (2024)
- Sun et al. MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices (2020)
- MLC-LLM project
- llama.cpp project