Skip to content

Reading Paths

At a glance Inference papers are vast and deep, and diving in blindly only leads to burnout. This page gives three paths — the engineer-onboarding line, the systematic deep-dive line, and the researcher line — plus the three-pass method, so you can pick a route by goal and read at a steady rhythm.

Reading Paths ​

One-sentence positioning: reading papers is not about "reading many" but about "picking the right path" — choose your goal first, then decide which papers to read and how deeply. This page gives you three curated roadmaps plus a general method for reading any systems paper thoroughly.

1. Why Read by Path: Navigation for an Ocean of Papers ​

Start with a few numbers: in the three arXiv categories cs.LG (machine learning) + cs.DC (distributed computing) + cs.OS (operating systems) alone, about 30-50 new LLM-inference-related papers appear on arXiv every week in 2024 — FlashAttention variants, speculative-decoding improvements, new quantization formats, new scheduling strategies, new hardware adaptations. Counting the backlog, papers in the inference acceleration field since 2017 (the year the Transformer was born) have long exceeded ten thousand.

At this scale, two brutal facts hold:

  1. No one can read them all. Even first-line engine developers closely read at most two or three papers a week and skim a dozen; the rest are filtered out by title, author, and second-order triage.
  2. Not filtering is a disaster. The quality distribution is extremely uneven: there are founding works that changed entire engine designs (FlashAttention, PagedAttention), engineering tuning reports ("we ran FA2 + AWQ on Llama-3 and got X throughput"), and padding work published just to be published. Read indiscriminately and, three months later, you will hold a pile of kernel-technique fragments while being unable to state the skeleton of the field.

This is why "paths" exist. Path = goal + order + depth standard:

text
No path:   scroll arXiv every week -> read a new quantization paper today -> a new scheduling one tomorrow
           -> three months later: seen everything, can explain nothing

With a path: set a goal first (engineering deployment / systematic deep dive / research)
           -> pick a route (onboarding / deep dive / research)
           -> read in the curated order (each paper builds on the previous one)
           -> three months later: hold a structured map

Follow a path and every paper you read is no longer an isolated file but a node on a map — the bottleneck raised by the previous paper is exactly the one answered by the next. This interlocking of knowledge nodes compounds your understanding.

How the path relates to the map, close-reading, and frontier

This papers section is a set of companions: Paper Map answers "what key papers exist in this decade and in what order" (the macro coordinate system); Classic Papers in Depth answers "how to master a single paper" (the microscope); Frontier Advances answers "where the field is heading" (the radar). The Reading Paths on this page turns those three into an action plan: read in order, trade off by goal, persist at a rhythm.

2. Overview of the Three Paths ​

By a learner's core goal, this site divides paper reading into three paths. They are not "levels" but "directions":

PathCore goalTime budgetWho it fitsEnd output
① Engineer OnboardingBuild the confidence and global view that "I can read a systems paper"~6 hoursLearners who just finished What Is Inference Acceleration? and want to touch inference papers for the first timeRead 6 papers thoroughly; can explain "why HBM is the inference bottleneck and why PagedAttention was revolutionary"
② Systematic Deep DiveMaster the full skeleton of inference acceleration~3 weeks (one group per week)Job hunters, engine engineers, master's/PhD students planning to go deep on systemsThe mechanisms of four topic groups + ablation logic + interview material
③ Researcher TrackKeep pace with the latest 2024-2026 trends and produce outputOngoing (4-6 hours per week)Advanced learners already in industry or academiaA stable information pipeline: can find, filter, judge, and reproduce
text
Engineer onboarding ── 6 hours ──► read: build confidence and a global view (first contact with inference papers)
Systematic deep dive ── 3 weeks ──► master: full skeleton + mechanisms + talking points (job hunting / leveling up)
Researcher track    ── ongoing ──► keep pace: find, filter, reproduce, produce (research / systems)

The three paths are not either/or — they are three stages

The engineer-onboarding line is the starting point for everyone, including those who will later go into research. Spend 6 hours building the confidence that "I can read this," then decide whether to invest 3 weeks in a systematic deep dive, and finally turn the researcher track into a daily habit. Skipping the first stage to chew the latest arXiv papers directly is why most people burn out on inference papers — the IO-complexity derivation in FlashAttention makes anyone who has not read FA1 give up on the spot.

3. Path 1: Engineer Onboarding (~6 Hours) ​

3.1 Route Design: Why These Six Papers ​

The onboarding line has only six papers, and the order is deliberately not chronological — it follows "cognitive dependencies": first use FlashAttention to build the intuition that "operators are the bottleneck," then add the systems perspective (PagedAttention) and the compression perspective (GPTQ/AWQ/EAGLE).

text
FlashAttention v1 (2022) -- operator bottleneck: the HBM bandwidth wall + kernel fusion
      |
FlashAttention v2 (2023) -- operator evolution: parallelism and work partitioning
      |
PagedAttention / vLLM (2023) -- system bottleneck: KV cache fragmentation + throughput
      |
GPTQ (2022) -- quantization: protecting important weights with Hessian information
      |
AWQ (2023) -- quantization: salient weights + activation awareness
      |
EAGLE-2 (2024) -- speculative decoding: a dynamic draft tree lets a small model guess what is next

3.2 One Line per Paper ​

OrderPaperOne-line why read it
1FlashAttention v1 (Dao et al., NeurIPS 2022)It shows how "operator IO-awareness" makes attention 2-4x faster without changing the math, building the core intuition that "the inference bottleneck is IO, not FLOPs"
2FlashAttention-2 (Dao, 2023)It rebuilds parallelism and work partitioning on v1, reaching 50-73% of theoretical peak on A100 — the best sequel to v1
3PagedAttention / vLLM (Kwon et al., SOSP 2023)It manages the KV cache like virtual memory, lifting LLM serving throughput 2-4x — the revolutionary starting point of LLM serving
4GPTQ (Frantar et al., ICLR 2023)It uses Hessian information for post-training 4-bit quantization, showing how the "protect important weights" paradigm keeps accuracy on 175B models
5AWQ (Lin et al., MLSys 2024)It finds that "salient weights" can be located via activation statistics — faster than GPTQ with comparable accuracy, the common production choice
6EAGLE-2 (Li et al., EMNLP 2024)It uses a dynamic draft tree to push the speculative-decoding acceptance rate from 50% to 70%+ — the best sequel to Medusa/EAGLE-1

3.3 Execution Cadence ​

  • Hour 1: read FlashAttention v1. Focus on the abstract, the Section 3 algorithm (the loop unrolling in Figure 1), and the Section 5 experiments. Pair it with Tri Dao's blog or an illustrated explainer to build the "IO-aware" intuition.
  • Hour 2: read FlashAttention-2. Focus on Section 3, "work partitioning" — why could v1 not sustain long sequences, and how does v2 split the work? Compare it against the FlashInfer blog.
  • Hour 3: read the vLLM paper. Focus on Section 3, "PagedAttention" (the block table in Figure 3) and Section 4, "continuous batching." Key intuition: the KV cache is the "virtual memory" of LLM serving, and fragmentation is the throughput killer.
  • Hour 4: read the GPTQ paper. Focus on Section 2, "OBQ and quantization order," and the Section 4 experiments. Key intuition: quantization is not uniform compression but differentiated protection by importance.
  • Hour 5: read the AWQ paper. Focus on Section 3, "salient weights," and Section 4, "activation awareness." Key intuition: activation statistics locate the important weight columns faster than a Hessian.
  • Hour 6: read the EAGLE-2 paper. Focus on Section 3, "dynamic draft tree," and the Section 4 experiments. Key intuition: speculative decoding trades a small model for the large model's compute, and the shape of the draft tree decides the acceptance rate.

Acceptance criteria for the engineer-onboarding line

After the six papers, you should be able to explain to someone without notes three things: (1) the inference bottleneck is IO, not FLOPs (FA); (2) LLM serving throughput depends on KV cache management, not model size (vLLM); (3) quantization and speculative decoding are two routes to "compressing cost" — the former compresses weight size, the latter compresses the number of decode steps. If you can state these three, your first attempt at reading inference papers has succeeded.

The deeper version of these six papers

The onboarding line is a "tasting" — read only the abstract, introduction, key figures, and conclusions, about 6 hours total. When you decide to move them into the Systematic Deep Dive, use the three-pass method from Section 6 to master each one; Classic Papers in Depth already has section-by-section breakdowns of FlashAttention v1/v2, vLLM, GPTQ, AWQ, and EAGLE-2 you can read side by side.

4. Path 2: Systematic Deep Dive (~3 Weeks) ​

4.1 Route Design: Four Topic Groups ​

The systematic deep-dive line divides inference acceleration into four topic groups, one per week, 2-3 papers per group, about 3 weeks total. The grouping follows the natural order of the field's evolution: operators first (FA), then systems (vLLM), then compression (quantization/speculation), and finally new paradigms (prefill/decode disaggregation + new operator formats).

text
Week 1  Operator layer        FA v1 -> FA v2 -> FlashInfer    "kernel fusion and IO-awareness"
Week 2  System layer          vLLM -> Orca -> SARATHI         "batching and scheduling"
Week 3  Compression layer     GPTQ -> AWQ -> SmoothQuant      "quantization and weight protection"
Week 4  Decode + disaggregation Medusa -> EAGLE-2 -> DistServe "speculative decoding and pipeline disaggregation"

4.2 Group Plans and Checkpoints ​

Week 1: Operator layer (FA v1 → FA v2 → FlashInfer)

PaperYearCore contribution
FlashAttention v12022IO-aware attention operator that moves the intermediate softmax state from HBM to SRAM, 2-4x faster on A100
FlashAttention v22023Rebuilt parallelism and work partitioning, raising GPU utilization on long sequences from 35% to 50-73%
FlashInfer2024Block-sparse + composable KV formats, unifying paged KV and variable-length attention in one kernel
  • What to grasp this week: the bandwidth gap between HBM and SRAM (~30x); why softmax must be computed in blocks; why recomputation beats storing the intermediate state.
  • Checkpoint: sketch FlashAttention's two-level structure from memory — "outer loop over K/V blocks, inner loop over Q blocks" — and explain why it cannot be inverted.
  • Operator mechanics: see Kernel Fusion and Custom Kernels and GPU Architecture and Optimization.

Week 2: System layer (vLLM → Orca → SARATHI)

PaperYearCore contribution
Orca2022Iteration-level scheduling that lets requests of different lengths advance in the same batch — the predecessor of continuous batching
vLLM / PagedAttention2023Manage the KV cache like virtual memory with a block table and block-level allocation, lifting LLM serving throughput 2-4x
SARATHI2023Chunked prefill: split a long prefill into chunks that share a batch with decode, balancing compute and memory
  • What to grasp this week: how the physical layout of the KV cache decides throughput; why request-level batching is inefficient for LLMs; how prefill and decode have exactly opposite compute/memory profiles.
  • Checkpoint: draw vLLM's block table and the logical view of PagedAttention, and explain what leftover problem of vLLM SARATHI solves.
  • Serving mechanics: see Batching and Request Scheduling and Model Serving and Orchestration.

Week 3: Compression layer (GPTQ → AWQ → SmoothQuant)

PaperYearCore contribution
GPTQ2022Second-order (Hessian) post-training quantization, keeping a 175B model's accuracy loss under 1% at 4-bit
AWQ2023Activation-aware salient-weight protection — 3x faster than GPTQ with comparable accuracy, a common production choice
SmoothQuant2022Smooth activation outliers into the weights, making W8A8 quantization of both weights and activations possible
  • What to grasp this week: why 4-bit weights are the sweet spot; why activation quantization is harder than weight quantization (outliers); how smoothing is essentially "difficulty migration."
  • Checkpoint: state the core difference between GPTQ and AWQ ("uses Hessian information" vs "uses activation statistics") and point out whether SmoothQuant solves the weight side or the activation side.
  • Quantization mechanics: see Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

Week 4: Decode layer + pipeline disaggregation (Medusa → EAGLE-2 → DistServe)

PaperYearCore contribution
Medusa2024Multi-head parallel prediction of the next k tokens — upgrading speculative decoding from "a small model guesses" to "the main model's extra heads guess"
EAGLE-22024Dynamic draft tree that pushes the acceptance rate from EAGLE-1's 50% to 70%+
DistServe2024Physically separate prefill and decode onto different GPUs, each scaling independently
  • What to grasp this week: the essence of speculative decoding is "small cost for big steps"; the draft tree's shape decides the acceptance rate; why the opposite compute/memory ratios of prefill and decode make disaggregation beneficial.
  • Checkpoint: draw the structural difference between the Medusa and EAGLE-2 draft schemes, and explain why DistServe is steadier than vLLM on tail latency.
  • Speculative-decoding cases: see Speculative Decoding and Medusa/EAGLE and Distributed Inference (TP/PP).

4.3 Execution Discipline for the Systematic Deep Dive ​

  • Fix 1-2 hours every day; do not cram. Leave a day between groups to "close the laptop and retell."
  • Every paper must produce notes: a three-sentence summary + one mechanism diagram + one "what happens if this design is removed" counter-question. The note method is in Reading Discipline & FAQ.
  • After each group, run a review: what is the internal evolution logic of this topic? What leftover problem of the previous paper does the next one solve?

Two hurdles you may hit on the deep dive

First, CUDA complexity discouragement — it is easy to get stuck when you reach FlashAttention's "online softmax + blocked loops." The countermeasure: intuition first (what is this paper doing), then pseudocode (the loop order in Algorithm 1), and only then the GPU thread mapping; most readers need go no further than the second level. Second, the reproduction temptation — wanting to spin up vLLM immediately to compare against a baseline. Countermeasure: the deep-dive stage is about understanding; leave hands-on reproduction for afterward; for the engine-comparison route see Inference Engine Comparison and Inference Benchmarking in Practice.

5. Path 3: The Researcher Track (Ongoing) ​

The systematic deep dive gives you the "stock" skeleton; the researcher track handles the "increment" — continuously wiring new papers into your knowledge map and producing your own work. The researcher track is not "refreshing the feed all day" (that only manufactures anxiety) but a loop of pipeline + reproduction + output:

text
Sources (arXiv / top conferences / engineering blogs / repo PRs)
      | filter
Candidate list (10-20 papers per week)
      | rank
Worth reading (3-5 papers per week) -> skim 2 / deep-read 1
      | reproduce
Pick 1 for a mini-reproduction (a kernel or scheduling snippet)
      | output
One-line card into the library + weekly notes + monthly review

5.1 Primary Sources: arXiv and Systems Conferences ​

arXiv is where every inference paper is published first. How to use it:

  • Subscribe by category: cs.LG (machine learning), cs.DC (distributed), cs.OS (operating systems), cs.PF (performance). When the daily update lands, scan titles only, and open the abstract only when a title makes you think "I should take a look."
  • Follow specific authors and labs: Tri Dao (Princeton/Together), Lianmin Zheng (Berkeley/DeepSeek), Ying Sheng (MIT/Together), Zhuoming Chen (Princeton), the DeepSeek AI systems group, NVIDIA's TensorRT-LLM team, the vLLM core team. Bookmark their pages and check their latest submissions periodically — far more efficient than scanning an entire category every day.
  • Use citation tracking: when you find a good paper, look at its "cited by" follow-ups on Google Scholar — it automatically walks you forward along the timeline.

5.2 Repos and PRs: The Second Pipeline of Engineering Signals ​

Inference deployment has a distinctive phenomenon: much important work appears in GitHub PRs before it becomes a paper. Repos worth watching long-term:

RepoValue
vllm-project/vllmMainline PRs hide the evolution of PagedAttention, the landing of chunked prefill, and the integration of various speculative-decoding schemes
NVIDIA/TensorRT-LLMFirst-hand signal of engineering-hardware co-design; FP8/FP4 kernels usually land here first
flash-attentionThe official implementation of the FA family and related operators; issues contain many edge cases
flashinfer-ai/flashinferThe official implementation of composable KV formats and block-sparse attention
sgl-project/sglangThe engineering frontier of RadixAttention and speculative decoding

5.3 Top Conferences: How the "Big Four" Systems Venues Divide the Work ​

ConferenceFull nameFocusScheduleWebsite
SOSP/OSDISymposium on Operating Systems / Operating Systems Design and ImplementationThe top systems conferences; serving-system papers such as vLLM, Orca, and AlpaServe often appear hereAlternate every Oct-Novsigops.org
MLSysConference on Machine Learning and SystemsThe dedicated ML-systems conference; FA, AWQ, and SARATHI often appear hereEvery Marchmlsys.org
ASPLOSArchitectural Support for Programming Languages and Operating SystemsThe hardware-software interface; Splitwise- and vAttention-style workEvery March-Aprilasplos-conference.org
ISCA/HPCAInternational Symposium on Computer Architecture / High-Performance Computer ArchitectureHardware-oriented; Splitwise and Mooncake often appear at HPCAEvery June / Apriliscaconf.org / hpca-conf.org

Practical advice: rather than refreshing arXiv and hoping, watch three things each year — the SOSP/OSDI accepted-papers list (find serving-system trends), the MLSys best papers (find operator and system hot spots), and the ASPLOS GPU-related papers (find hardware co-design). Skim titles and abstracts for two hours each time.

5.4 Rhythm of the Researcher Track ​

  • 4-6 hours per week: scan arXiv + repo PRs -> pick 10 papers to read abstracts -> pick 2-3 to skim -> pick 1 to deep-read with the three-pass method + a mini-reproduction.
  • One review per month: what were the main trends this month? Which are engineering optimizations and which are genuine algorithmic innovation? Which will enter mainstream engines and which will disappear?
  • Notes and FAQ for the researcher track are in Reading Discipline & FAQ; this site's Frontier Advances has already systematically sorted the important 2024-2026 breakthroughs (FP8/FP4, prefill/decode disaggregation, EAGLE-3, new hardware) and can be your starting point for tracking.

The right mindset for tracking the frontier

The value of the researcher track is not "being first to know" but "knowing more accurately than others." Reading a paper two weeks late is perfectly fine; judging it more reliably than the retellers (how was this speedup obtained, what is the cost, where are the limits) is the advanced learner's real skill. Slow and accurate beats fast and panicked.

6. How to Read a Paper: The Three-Pass Method ​

Whichever path you take, when it comes down to "reading one paper," use the three-pass method proposed by computer scientist S. Keshav — each pass goes deeper, and each pass allows "early termination":

PassTimeWhat to readOutputWhen you can stop here
Pass 1: bird's-eye5-10 minTitle, abstract, introduction, section headings, all figures, conclusionCan answer "what does this paper do and what is its contribution"It is outside your interest/direction — discard it
Pass 2: close read~1-2 hoursMethod, experiments, every figure detail; mark key definitions and pseudocodeCan retell the paper's full argument to someone elseYou only need to "use" it, not "modify" it
Pass 3: rebuildHours to daysLine-by-line reading, re-derive, reimplement the kernel, question every assumptionCan point out the paper's holes and room for improvementYou are doing research, reproduction, or a survey
text
Pass 1  5-10 min   -> abstract + intro + figures + conclusion  -> decide: read or not
Pass 2  1-2 hours  -> method + experiments, mark the pseudocode -> decide: how deep to use it
Pass 3  hours+     -> reimplement the kernel, challenge assumptions -> decide: follow up or not

For inference-systems papers, the three-pass method has three extra points:

  1. Always read abstract -> profile figures -> conclusion first. A systems paper's "profile figures" are extremely dense: one roofline figure is worth three paragraphs, and one batch-throughput curve exposes the method's true level. Look at the figures first, then read the text with questions in hand.
  2. In Pass 2, always do "ablation tracking." Wherever the paper shows "Table X: Ablation," stop and ask: what did the authors prove by removing this component? This is the core method for distinguishing "genuine innovation" from "engineering stacking." FlashAttention doubling its numbers without an IO-awareness ablation is far less credible than "remove IO-awareness and the speed drops to 1.1x."
  3. In Pass 3, translate every "speedup" into a "boundary." Attach a qualifying condition to every 2x and 3x: under what model, what batch, what sequence length, and what hardware does it hold? Once translated, you have truly read the paper (more methodology in Reading Discipline & FAQ).

7. Trade-offs: Deep Read vs. Skim ​

The hardest resource constraint in reading papers is time, so every strategy revolves around one trade-off: how to split limited energy between "deep reading" and "skimming."

DimensionSkimDeep read
Cost per paper10-30 min4-10 hours
What you gainBreadth: know "what exists, where the trend is"Depth: know "why, and where the boundary is"
Applies toTracking the frontier, screening candidates, building a global viewPapers strongly tied to your work/research, classic landmarks, papers you will reproduce
RiskSee the forest but not the trees; parroting othersSee the trees but not the forest; a narrowing view

In practice, use the 80/20 rule: dispose of 80% of papers with Pass 1 (bird's-eye), 15% with Pass 2 (close read), and only 5% deserve Pass 3 (rebuild). Only three kinds of papers are truly worth a deep read:

  1. Paradigm-level classics: papers that changed the field's thinking (e.g. FlashAttention, vLLM) — deep-read them, because they are the cornerstone of everything you understand later.
  2. Papers strongly tied to your current work/research: papers that connect directly to your deployment scenario — deep-read them, study their ablations and limits, and judge whether they port to your setting.
  3. Papers that are heavily cited yet contested: unavoidable and controversial — deep-read them and decide with your own judgment.

Everything else (hot news, marginal improvements, pure engineering reports) gets a skim. Deep reading is an investment; skimming is overhead — use overhead to maintain breadth and investment to buy depth; that is the most sustainable reading strategy in this field.

How to pair this with the rest of the site

While reading a path, treat the other sections as "accelerators": Paper Map helps you locate a paper on the map (source / milestone / survey) and decide skim vs. deep read; Classic Papers in Depth has done Pass-3-level breakdowns of the most important papers for you; Frontier Advances slots the newest papers onto the trend line; Reading Discipline & FAQ handles "can't understand, can't remember, can't keep going."

8. Summary: Choosing Your Path ​

  • New to inference deployment, reading papers for the first time -> take the Engineer Onboarding line: 6 papers in 6 hours to build confidence.
  • Need to interview / ship / build an engine -> take the Systematic Deep Dive line: 12 papers in 3 weeks to build the skeleton.
  • Want to do research / write papers / do open source -> after the deep dive, switch to the Researcher Track and turn "pipeline + mini-reproduction" into a habit.
  • Not sure which one -> start with the Engineer Onboarding line; after 6 hours you will know where to go.

9. Further Reading ​

You now have three paths and the three-pass method. Pick the next stop by your goal:

  • Start Here — the entry to the papers section: three values of reading papers and the single most important advice
  • Paper Map — the macro coordinate system of the deep-dive line: key papers of 2014-2025 by topic and timeline
  • Classic Papers in Depth — section-by-section breakdowns of the classics touched by the onboarding and deep-dive lines
  • Frontier Advances — the starting point of the researcher track: a systematic sort of the important 2024-2026 breakthroughs
  • Reading Discipline & FAQ — a companion to the three-pass method: note-taking, what to do when stuck, how to judge a paper's quality
  • What Is Inference Acceleration? — from papers back to the system: the basics of inference deployment
  • Learning Paths: Three Routes — the site-wide knowledge learning route
  • Curated Resources — a real collection of paper code, benchmarks, and tools

References ​

The following are all real public resources for self-study: