Skip to content

Start Here

At a glance Why should you read papers even if you work on inference acceleration and deployment? This page explains the distinct value of three kinds of papers, offers three entry routes by goal, introduces the five-part navigation of this papers section, and closes with the single most important piece of advice for reading papers.

Start Here ​

One-sentence positioning: this section teaches you to read the papers of the inference acceleration and deployment field — from "can parse a SOSP/MLSys/arXiv abstract" to "can keep pace with the systems rhythm of the H100/B200 era." It does not ask you to swallow the entire FlashAttention derivation in one go; it first solves three problems: what to read, how to read, and how deeply to read.

1. What This Section Is: Why Read Papers for Inference Acceleration ​

Reading papers is not the patent of researchers, nor a "bonus performance" to recite in interviews. For anyone serious about inference deployment, it has three irreplaceable values:

1.1 Understand the Principles, Not the Configurations ​

You will learn engine.generate(...) from framework docs, but only a paper can tell you why the design behind that line exists. Why does PagedAttention manage the KV cache with a block table? Why does FlashAttention prefer more compute to fewer HBM round-trips? Why does GPTQ do column-wise merging on the Hessian instead of per-layer clipping?

None of these were pulled from thin air. Every design corresponds to a failure, a profile, or an ablation in the original paper. What you get from a paper is not "how to tune the parameters" but "why it was designed this way" — the latter is the watershed between "someone who can use a framework" and "someone who can write an engine." The core knowledge on this site (such as What Is Inference Acceleration?) gives you the conceptual map; the papers section fills in why every place on that map is named what it is.

1.2 Keep Up with the Frontier, Not the Hype ​

Inference systems iterate on a quarterly clock: what you learned as "the latest" a year ago may already be in a second-tier engine's changelog. Papers are the first-hand source of frontier information — new operators, new schedulers, and new quantization formats always appear as papers first, and only reach the main branches of vLLM, TensorRT-LLM, and SGLang six months later.

People who do not read papers get filtered, simplified, or even missed information from release notes; people who read papers can directly judge "on what model, what batch, what hardware, and under what metric was this 3x speedup achieved." The former chase changelogs; the latter watch the trend line.

1.3 A Hard Credential — and an Irreplaceable One ​

In inference-deployment interviews, "talk through a systems paper" appears almost without fail. What interviewers ask is usually not the abstract, but:

  • What bottleneck does this paper solve? Why is it the bottleneck?
  • What is its core contribution? Why did previous methods not work?
  • What do the ablations prove? Which parts are genuine algorithmic innovation, and which are just kernel tuning?
  • Transplanted to my hardware (A100, not H100) and my model (13B, not 70B), where would this paper's method break?

Only those who have closely read the original can answer these. What the interviewer wants is not "you have memorized PagedAttention" but "you possess the ability to read systems papers" — because that is precisely the core skill of daily deployment optimization.

2. Where to Start: Three Routes by Goal ​

There is no single standard for how deeply to read a paper; pick a route by your goal. All three routes are expanded in Reading Paths; here is a quick overview first:

text
Goal: two-hour onboarding
├─ Read only FlashAttention v1 (abstract + Section 3 algorithm + Figure 1)
├─ Add one illustrated explainer or a set of notes
└─ Output: can explain "why IO-awareness makes attention 2-4x faster"

Goal: engineering deployment
├─ Deep-read vLLM / FlashAttention / GPTQ / AWQ / EAGLE-2 in full
├─ Focus on the profiles, the ablations, and the engineering limits (Discussion)
└─ Output: can judge "can this method be ported to my business model and batch size"

Goal: research
├─ Deep-read everything + math derivation + reimplement the kernel yourself
├─ Read the references backward and compare follow-ups (FA1 -> FA2 -> FA3)
└─ Output: can propose your own operator/scheduling improvement and benchmark it
RouteWho it fitsReading depthPapers to read
Two-hour onboardingLearners who just finished What Is Inference Acceleration?Abstract + introduction + algorithm figure + conclusionStart with the FlashAttention entry in Classic Papers in Depth
Engineering deploymentEngineers who want papers to solve real deployment problemsFull text + ablations + limitationsCombine Classic Papers in Depth with Frontier Advances
ResearchSystems-track graduate students planning to publish at MLSys/SOSPFull text + derivations + kernel reimplementationBuild a global coordinate system from the Paper Map, then go deep

One sentence for newcomers

For your first paper, always pick a "landmark and easy to read" one, not the "newest and hottest." FlashAttention v1 (NeurIPS 2022) has stood the test of time, is cited by countless follow-ups, and is clearly written. First build the confidence that "I can read a systems paper," then worry about quantity.

3. What Is in This Section ​

This papers section has five parts; visit them in this order:

PageOne-line introWhat you get
Reading PathsThe section's main entry: specific paper lists and reading orders for the three routesA paper-reading roadmap arranged by goal
Paper MapMaps the key inference-acceleration papers of 2014-2025 by topic and timelineA macro coordinate system: which paper is the source, which is the milestone, which is the survey
Classic Papers in DepthClose-reading of the ten papers that changed the field: background, mechanism, experiments, limitsThe ability to master principles, plus interview material
Frontier AdvancesTracks recent breakthroughs and trends: FP8/FP4, prefill/decode disaggregation, speculative decoding, new hardwarePace with the 2024-2026 frontier
Reading Discipline & FAQMethodology for reading papers: how to take notes, what to do when stuck, how to judge a paper's qualityA sustainable reading habit, not a one-off impulse

The logic across the section is map first, then close-reading, then frontier: the map answers "what to read," close-reading answers "how deeply," and the frontier answers "where to go next." Consult the Glossary whenever a term is unfamiliar.

4. The Single Most Important Piece of Advice ​

The most important advice: take away the mechanism, not the number

What sticks most easily from a paper is often the speedup figure: "2-4x faster than PyTorch," "2-3x the throughput of vLLM," "EAGLE-3 with a 6.5x speedup on MT-Bench." But remember:

Every speedup is the product of four things — model + batch + hardware + evaluation protocol. Change the model (13B instead of 70B), the hardware (A100 instead of H100), the batch size (1 instead of 256), or the evaluation script (vLLM 0.6 instead of 0.4), and the number can collapse entirely. Reproduce a 2022 FlashAttention v1 experiment on a 2024 H100 and the conclusion might even flip.

What is truly worth taking away are three things:

  1. The mechanism — why does this method work? What structural bottleneck does it solve? (Example: FlashAttention moves the softmax intermediate state from HBM to SRAM, defeating the HBM bandwidth wall.)
  2. The failure analysis — which experiments did the authors skip, on which hardware does it not apply, what limitations do the authors themselves state? Failures are often worth more than successes.
  3. The boundary of applicability — under what batch, what sequence length, and what hardware generation does it hold? (Example: PagedAttention pays off most at long sequences and large batches, and barely helps a single user on short sequences.)

Benchmark numbers expire; mechanisms and boundaries do not.

Further Reading ​

References ​

The following are all real public resources for self-study:

  • arXiv preprint server — where most inference-systems papers are released as preprints; focus on cs.LG, cs.DC, cs.OS
  • Papers with Code — papers + code + inference-benchmark aggregation; handy for checking SOTA and reproductions
  • FlashAttention official repo — the companion code of the paper; the best crutch for reading the FlashAttention series
  • vLLM official repo — the industrial-strength implementation of PagedAttention; issues and PRs hide a wealth of engineering detail
  • SOSP 2023 proceedings — the conference where the vLLM paper appeared; read DistServe and other systems papers from the same year
  • CUDA Programming Guide — the hardware background required for reading operator papers (FlashAttention, FlashInfer)
  • Tri Dao's blog — the FlashAttention author's blog, often with "author's-view" follow-up notes