Skip to content

Paper Map

At a glance A searchable map of the key inference-acceleration and deployment papers from 2014-2025, grouped by theme and covering six areas — operator optimization, quantization, serving systems, speculative decoding, distributed training/inference, and on-device deployment — with a paradigm-shift timeline and a task-based paper index.

Paper Map ​

One-sentence positioning: this map lays the key inference-acceleration and deployment papers of 2014-2025 along the two axes of theme and year, forming a coordinate system you can search at any time — from PipeDream's pipeline parallelism in 2014 to EAGLE-3's 6.5x speedup in 2025. After reading it, you can state "what bottleneck each key paper solved, which papers it connects to before and after, and whether it deserves a close read or a skim."

The biggest obstacle to reading papers is not "failing to understand one" but "not knowing which one to read next." Tutorials teach in order, but the world of papers is organized by problems: the answer to the same problem ("how to keep attention from being stuck on HBM") was FlashAttention v1 in 2022, FlashAttention-2 in 2023, FlashAttention-3 in 2024 (H100 FP8), and also vAttention in 2024 (using software PagedAttention to approximate hardware-native sparse attention). If you read only along the timeline, you get lost in "who optimized whom"; if you read only by theme, you miss the macro narrative of "why hardware-generation shifts reshape algorithmic routes." The purpose of this map is to overlay theme and time so that, facing any paper, you can answer three questions: which technical line does it belong to? where does it sit in the paradigm evolution? and which papers form a "must-read chain" with it?

How the map divides work with the other pages

This papers section has five pages with different roles: Start Here explains "why read inference papers," Reading Paths explains "how to read by goal," this page (Paper Map) explains "what papers in the field are worth reading and how they relate," Classic Papers in Depth digs into papers one by one, and Frontier Advances tracks the new trends of 2024-2026. In one sentence: the map sets coordinates, close-reading gives depth, and the frontier points the direction.

1. How to Read This Map ​

Before unfolding the map, three conventions:

  1. Numbering convention: every paper is given an arXiv number (or the formal venue), so the full text is free at https://arxiv.org/abs/NUMBER; older papers without an arXiv entry are marked "—" with the venue given, and the original links are in the references.
  2. Year convention: a paper's year always means its first public release (formal venue publication, or arXiv preprint submission). For example, the FlashAttention v1 preprint was submitted in June 2022 and it appeared at NeurIPS 2022, so the table records 2022.
  3. Deep-read convention: the map gives only a "one-line contribution" — an answer to "why is this paper worth remembering." To master a paper, jump to Classic Papers in Depth; to chase its follow-ups, use Frontier Advances and the search method in FAQ.

Use the three coordinates together: theme decides which line it belongs to, year decides where it sits in the paradigm evolution, and importance decides how much time you invest. Here is the overview of the map:

text
                    Theme dimension (six technical lines)
                    ┌────────────────────────────────────┐
                    │ 1. Operator optimization           │
                    │ 2. Quantization & compression      │
Time dimension |    │ 3. Serving systems & scheduling    │
                    │ 4. Speculative & parallel decoding │
                    │ 5. Distributed training/inference  │
                    │ 6. On-device & edge deployment     │
                    └────────────────────────────────────┘

2. The Academic Map by Theme ​

2.1 Operator Optimization: From Naive Attention to IO-Awareness and Hardware Co-Design (2017-2025) ​

The core question of this group is always: why is attention's O(n²) complexity the bottleneck — is it a FLOPs bottleneck or an IO bottleneck? After the Transformer in 2017, the community first worked on FLOPs (sparse attention, Longformer, Big Bird) until FlashAttention pointed out in 2022 that the bottleneck was HBM bandwidth, correcting the route back to "IO-awareness + operator fusion."

PaperAuthors/InstitutionYearVenue/arXivOne-line contribution
Attention Is All You NeedVaswani et al. (Google Brain)2017arXiv:1706.03762Proposes the Transformer and the pure-attention architecture — the source of O(n²) complexity
Longformer / Big BirdBeltagy et al. / Zaheer et al.2020arXiv:2004.05150 / arXiv:2007.14062Sparse attention replacing global attention with "local window + a few global tokens," O(n) compute on long documents
FlashAttention v1Dao et al. (Stanford)2022arXiv:2205.14135IO-aware attention operator: tiling + online softmax + no write-back of intermediates to HBM, 2-4x faster on A100
FlashAttention-2Dao (Stanford)2023arXiv:2307.08691Rebuilt parallelism and work partitioning, raising GPU utilization on long sequences from 35% to 50-73%
FlashAttention-3Shah et al. (Princeton)2024arXiv:2407.08608FP8 + asynchrony (warp-specialized) on H100, 1.2 PFLOPs peak, 1.5-2x faster than FA2
FlashInferYe et al. (Princeton/DeepSeek)2024arXiv:2401.12448Block-sparse + composable KV formats, unifying paged KV and variable-length attention in one kernel
vAttentionPrabhu et al. (Georgia Tech)2024arXiv:2405.04447Achieves software PagedAttention's function with hardware-native sparse attention, avoiding SM stalls

Why read the FlashAttention series before vAttention

The three FlashAttention generations (v1→v2→v3) show the evolution paradigm of "the same math formula, three generations of engineering": v1 found the bottleneck (IO), v2 optimized parallelism, v3 co-designed with hardware (H100 FP8). vAttention asks "can hardware sparse attention replace software PagedAttention," a new-paradigm debate from 2024. Read this line and you grasp the core law of "how operator optimization follows hardware generations." Operator mechanics in detail: Kernel Fusion and Custom Kernels.

2.2 Quantization and Compression: From Uniform Quantization to Second-Order Information and Rotation (2020-2025) ​

The core question of quantization: can weights and activations be represented with fewer bits at acceptable accuracy loss? This line went through three paradigm shifts in the 2020s: naive round-to-nearest → LLM.int8() discovers outliers → GPTQ/AWQ/SmoothQuant protect important parameters more cleverly → rotation quantization (QuaRot/SpinQuant) and FP8/FP4 hardware quantization in 2024.

PaperAuthors/InstitutionYearVenue/arXivOne-line contribution
LLM.int8()Dettmers et al. (Washington)2022arXiv:2208.07339Discovers "outlier features" in models >6.7B and proposes mixed precision (outlier columns in FP16, the rest INT8) — the start of LLM quantization research
GPTQFrantar et al. (IST Austria)2022arXiv:2210.17323Post-training 4-bit weight quantization with Hessian information, <1% accuracy loss on 175B models — the second-order paradigm
SmoothQuantXiao et al. (MIT/ByteDance)2022arXiv:2211.10438Smooths activation outliers into the weights, making W8A8 quantization possible — the icebreaker of activation quantization
AWQLin et al. (MIT/Hugging Face)2023arXiv:2306.00978Activation-aware salient-weight protection, 3x faster than GPTQ with comparable accuracy — a common production choice
ZeroQuantYao et al. (Microsoft)2022arXiv:2206.01861Joint weight-and-activation quantization + per-block calibration, an early industrial post-training scheme
FP8 LLM Inference (NVIDIA)NVIDIA2023NVIDIA FP8 white paperH100 introduces FP8 Tensor Cores — a hardware-level sweet-spot format
QuaRotAshkboos et al. (ETH)2024arXiv:2404.00456Rotation quantization: random orthogonal transforms "flatten" outliers, greatly improving W4A4 accuracy
SpinQuantLiu et al. (Meta)2024arXiv:2405.16406Learned rotation matrices, better than QuaRot's random rotation, W4A4 approaching W4A16

Two threads of quantization

Weight side: GPTQ and AWQ use second-order/activation information to locate "salient weights" — "differentiated protection." Activation side: SmoothQuant migrates difficulty from activation to weight, and QuaRot/SpinQuant "flatten" the distribution with rotation. The two threads converge in 2024 at W4A4 + rotation, letting 4-bit inference approach 16-bit accuracy. Quantization mechanics in detail: Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

2.3 Serving Systems and Scheduling: From Naive Batching to Prefill/Decode Disaggregation (2022-2025) ​

The core question of serving systems: with many users and variable-length requests, how do you keep the GPU from idling, memory from fragmenting, and tail latency under control? Orca proposed iteration-level scheduling in 2022, vLLM managed the KV cache like virtual memory in 2023, DistServe/Splitwise physically disaggregated prefill and decode in 2024, and SARATHI balanced compute with chunked prefill.

PaperAuthors/InstitutionYearVenue/arXivOne-line contribution
OrcaYu et al. (Microsoft)2022OSDI 2022Iteration-level scheduling — requests of different lengths advance in the same batch, the predecessor of continuous batching
vLLM / PagedAttentionKwon et al. (Berkeley)2023SOSP 2023Manage the KV cache like virtual memory with a block table and block-level allocation, 2-4x throughput
AlpaServeLi et al. (Berkeley)2023OSDI 2023Model parallelism + time multiplexing, using statistical multiplexing to guide multi-model deployment
SARATHIAgarwal et al. (Microsoft)2023arXiv:2308.16369Chunked prefill: split a long prefill into chunks that share a batch with decode, balancing compute and memory
DistServeZhong et al. (Peking/Tsinghua)2024arXiv:2401.09670Physically separate prefill and decode onto different GPUs, each scaling independently
SplitwisePatel et al. (Microsoft/IIT)2024HPCA 2024Similar to DistServe but focused on cluster-level "phase-specific machine" scheduling
SGLangZheng et al. (Berkeley/DeepSeek)2024arXiv:2312.07104RadixAttention: a prefix tree reuses the KV cache — a throughput weapon for multi-turn dialogue and few-shot scenarios
MooncakeMoonshot AI2024arXiv:2407.00079A planet-scale serving architecture that disaggregates the KV cache, treating it as a first-class citizen

Three stages of serving-system evolution

Stage 1 (Orca, 2022): swap "batching by request" for "batching by iteration" — a throughput leap. Stage 2 (vLLM, 2023): swap "pre-allocated memory" for "on-demand block-level allocation" — KV fragmentation eliminated. Stage 3 (DistServe/SARATHI, 2024): disaggregate "prefill + decode on the same GPU in the same batch," because their compute/memory ratios are exactly opposite. Read these three stages and you understand the history of LLM serving. Serving mechanics: Batching and Request Scheduling and Model Serving and Orchestration.

2.4 Speculative and Parallel Decoding: From Single-Step to Multi-Token Prediction (2022-2025) ​

The core bottleneck of the decode layer: autoregressive generation produces one token per step, recomputing the KV cache 1000+ times just to generate 1000 tokens. Speculative decoding's idea is "guess k tokens at small cost, then verify them with the large model in one pass." Leviathan formalized it in 2023, Medusa/EAGLE-1/2 each upgraded drafting in 2024, and EAGLE-3 pushed the speedup to 6.5x in 2025.

PaperAuthors/InstitutionYearVenue/arXivOne-line contribution
Speculative DecodingLeviathan et al. (Google/Technion)2023arXiv:2211.17192A small model guesses + the large model verifies in parallel, lossless acceleration — the formalization of speculative decoding
SpecInferMiao et al. (CMU/SCU)2023arXiv:2302.02018Token-tree parallel verification + lion-hearted SSG, upgrading speculative decoding with multiple drafts
MedusaCai et al. (Princeton/DeepSeek)2024arXiv:2401.10774The main model predicts in parallel with multiple heads, no independent small model needed — simplifying speculative decoding
EAGLE-1Li et al. (Peking/Tencent)2024arXiv:2401.15077Drafts from the main model's hidden state, aligning draft and verify features, 50%+ acceptance rate
EAGLE-2Li et al. (Peking/Tencent)2024arXiv:2406.16858Dynamic draft tree that prunes the tree by context probability, pushing the acceptance rate to 70%+
EAGLE-3Li et al. (Peking/Tencent)2025arXiv:2503.01840Deeper hidden-layer features + an improved training objective, 3-6.5x speedup
SpecForgeSGLang team2025SGLang blogA multi-draft-model training pipeline, treating speculative decoding as an "engineering product" rather than a "model product"
DeepSeek-V3 MTPDeepSeek-AI2024arXiv:2412.19437The main model natively supports multi-token prediction, learning speculative-decoding ability at training time

The "cost ledger" of speculative decoding

The core trade-off of speculative decoding is: the extra compute of drafting vs. the steps saved by verification. The more accurate the draft and the higher the acceptance rate, the more is saved; the larger the draft, the higher the overhead. The evolution logic of the EAGLE series is "use the main model's hidden state to make the draft both accurate and small," while DeepSeek-V3's MTP takes another route — teach the main model to emit multiple tokens at once at training time, internalizing speculative decoding into the training objective. Which one becomes mainstream is an open question for 2025-2026. Cases: Speculative Decoding and Medusa/EAGLE.

2.5 Distributed Training/Inference: From Data Parallelism to Expert and Pipeline Parallelism (2019-2025) ​

Distributed is no longer "a training-time thing" — MoE inference makes expert parallelism a hard requirement for deployment, and DeepSeek-V3's 671B total / 37B active parameters make inference-side distributed scheduling as complex as training-side.

PaperAuthors/InstitutionYearVenue/arXivOne-line contribution
Megatron-LMShoeybi et al. (NVIDIA)2019arXiv:1904.10509Tensor parallelism: shard each layer's matrix multiplications by column/row across GPUs — infrastructure for training large models
PipeDreamNarayanan et al. (Stanford/Google)2019arXiv:1806.03312Pipeline parallelism + micro-batching + 1F1B scheduling, making "layer-sharding across GPUs" viable
MegaScaleJiang et al. (ByteDance)2024NSDI 2024An industrial system for trillion-parameter MoE training, with communication/compute overlap and fault recovery
Mixtral of ExpertsMistral AI2024arXiv:2401.04088Open-weights 8-expert MoE, bringing MoE from closed to open source — expert parallelism becomes standard on the inference side
DeepSeek-V3DeepSeek-AI2024arXiv:2412.19437256 experts + 8 active MoE + MTP, pushing open-source MoE inference serving to a hundred-billion-parameter benchmark

The training/inference boundary of distributed is blurring

Megatron and PipeDream are training-side papers, but their parallel schemes (TP/PP/SP) were inherited directly by inference — multi-GPU inference in TensorRT-LLM and vLLM uses the same sharding. MoE inference (Mixtral, DeepSeek-V3) brought "expert parallelism" from training into inference, forcing inference engines to add new scheduling. Today you cannot do inference deployment without understanding training-side parallelism. See Distributed Inference (TP/PP).

2.6 On-Device and Edge Deployment: Cramming an LLM into Phones and Laptops (2020-2025) ​

As LLMs move onto phones and laptops, the question shifts from "how to be faster" to "how to run at all in 4-8 GB of memory." This line includes on-device architectures (MobileBERT), on-device inference engines (MLC-LLM, llama.cpp), and on-device quantization (MCT-LLM).

PaperAuthors/InstitutionYearVenue/arXivOne-line contribution
MobileBERTSun et al. (HUST/Google)2020arXiv:2004.02584Distillation + bottleneck structures layer by layer, compressing BERT to run on a phone — the start of on-device Transformers
MCT-LLMLinux Foundation / ONNX Format2023Microsoft Build 2023Microsoft's on-device INT4 quantization, running Llama-7B on a Surface Laptop
MLC-LLMTVM team (OctoML/Apache)2023mlc.aiCompile LLMs to GPU/NPU/CPU with TVM — a compiler approach to on-device LLMs
llama.cppGerganov et al.2023github.com/ggerganov/llama.cppPure C++ + the GGUF quantization format, running Llama on CPUs and low-end GPUs — the on-device de facto standard

Two routes for on-device inference

The compiler route (MLC-LLM): treat an LLM as a computation graph and compile it to different hardware backends with TVM — elegant engineering but slow to adapt. The direct-implementation route (llama.cpp): rewrite the whole inference path in pure C++ + SIMD — flexible but adaptation is manual. Today llama.cpp is the on-device de facto standard thanks to its simplicity and active community, but whether the compiler route overtakes it in the NPU era is undecided. See Mobile Deployment and llama.cpp and GGUF.

3. Paradigm-Shift Timeline: A Decade of Evolution in One Diagram ​

Arrange the key papers of the six themes along a timeline and you can see the rhythm of paradigm shifts in inference acceleration:

text
2017  Transformer           -- the source of O(n^2) complexity
      |
2019  Megatron-LM/PipeDream -- parallel infrastructure for training large models
      |
2020  Longformer/BigBird     -- naive sparse attention (sidestepping O(n^2))
      |
2022  LLM.int8()             -- outliers discovered; quantization research breaks through
      |  GPTQ                -- second-order quantization
      |  SmoothQuant         -- activation smoothing
      |  Orca                -- iteration-level scheduling
      |  FlashAttention v1   -- IO-aware operator (paradigm turning point)
      |
2023  vLLM/PagedAttention    -- KV cache managed like virtual memory (paradigm turning point)
      |  AWQ                 -- activation-aware quantization
      |  FlashAttention-2    -- parallelism and work-partition optimization
      |  Speculative Decoding -- speculative decoding formalized
      |  llama.cpp/MLC-LLM   -- on-device LLM lands
      |  SARATHI             -- chunked prefill
      |
2024  FlashAttention-3       -- H100 FP8 co-design
      |  vAttention          -- hardware-native sparsity replaces software paged
      |  DistServe/Splitwise -- prefill/decode physically disaggregated
      |  SGLang              -- RadixAttention prefix-tree reuse
      |  Medusa/EAGLE-1/2    -- speculative decoding upgrades
      |  QuaRot/SpinQuant    -- rotation quantization
      |  Mixtral/DeepSeek-V3 -- MoE inference serving at scale
      |  Mooncake            -- KV-cache-disaggregated architecture
      |
2025  EAGLE-3                -- 6.5x speculative decoding
      |  FP4/NVFP4           -- Blackwell B200 era
      |  P-EAGLE/SpecForge   -- speculative decoding productionized
      |  ...

The two paradigm turning points are the most worth remembering: FlashAttention v1 in 2022 corrected the bottleneck from FLOPs back to IO, and vLLM in 2023 redefined LLM serving from "loading a model" to "managing memory." Prioritize close-reading papers from these two years, because they defined today's engineering coordinate system.

4. Task-Based Paper Index ​

Different roles and scenarios care about different papers. This table lets you look up papers by "task":

What you are doingMust-read papersAdvanced papers
Writing an attention kernelFlashAttention v1/v2FlashAttention-3, FlashInfer, vAttention
Optimizing LLM serving throughputvLLM, OrcaSARATHI, DistServe, Splitwise, SGLang, Mooncake
Quantizing a 70B model for launchGPTQ, AWQSmoothQuant, LLM.int8(), QuaRot, SpinQuant
Speeding up the decode phaseSpeculative Decoding, EAGLE-2Medusa, EAGLE-1/3, SpecInfer
Deploying an MoE modelMixtral, DeepSeek-V3MegaScale, Megatron-LM
Running on-device LLMsllama.cpp, MLC-LLMMobileBERT, MCT-LLM
Squeezing performance on H100/B200FlashAttention-3, FP8 inferenceNVFP4, QuaRot, vAttention
Cramming before an interviewAll of Classic Papers in DepthPair with Frontier Advances

5. Reading by Importance Tier ​

If time is short, pick papers by the tiers below:

Three-tier reading list

Must-read 5 (anyone doing inference deployment should read these)

  • FlashAttention v1 (2022)
  • vLLM / PagedAttention (2023)
  • GPTQ or AWQ (pick one)
  • EAGLE-2 (2024)
  • Orca or SARATHI (pick one)

Strongly recommended 5 (systematic deep dive or job hunting)

  • FlashAttention-2 (2023)
  • SmoothQuant (2022)
  • Medusa (2024)
  • DistServe (2024)
  • FlashAttention-3 (2024)

Researcher must-read 5 (frontier tracking / paper writing)

  • vAttention (2024)
  • SGLang (2024)
  • QuaRot or SpinQuant (2024)
  • DeepSeek-V3 Technical Report (2024)
  • EAGLE-3 (2025)

6. Further Reading ​

References ​

The following are all real public resources; the originals are directly accessible: