Appearance
Start Here
One-sentence positioning: this section teaches you to read the papers of the inference acceleration and deployment field — from "can parse a SOSP/MLSys/arXiv abstract" to "can keep pace with the systems rhythm of the H100/B200 era." It does not ask you to swallow the entire FlashAttention derivation in one go; it first solves three problems: what to read, how to read, and how deeply to read.
1. What This Section Is: Why Read Papers for Inference Acceleration
Reading papers is not the patent of researchers, nor a "bonus performance" to recite in interviews. For anyone serious about inference deployment, it has three irreplaceable values:
1.1 Understand the Principles, Not the Configurations
You will learn engine.generate(...) from framework docs, but only a paper can tell you why the design behind that line exists. Why does PagedAttention manage the KV cache with a block table? Why does FlashAttention prefer more compute to fewer HBM round-trips? Why does GPTQ do column-wise merging on the Hessian instead of per-layer clipping?
None of these were pulled from thin air. Every design corresponds to a failure, a profile, or an ablation in the original paper. What you get from a paper is not "how to tune the parameters" but "why it was designed this way" — the latter is the watershed between "someone who can use a framework" and "someone who can write an engine." The core knowledge on this site (such as What Is Inference Acceleration?) gives you the conceptual map; the papers section fills in why every place on that map is named what it is.
1.2 Keep Up with the Frontier, Not the Hype
Inference systems iterate on a quarterly clock: what you learned as "the latest" a year ago may already be in a second-tier engine's changelog. Papers are the first-hand source of frontier information — new operators, new schedulers, and new quantization formats always appear as papers first, and only reach the main branches of vLLM, TensorRT-LLM, and SGLang six months later.
People who do not read papers get filtered, simplified, or even missed information from release notes; people who read papers can directly judge "on what model, what batch, what hardware, and under what metric was this 3x speedup achieved." The former chase changelogs; the latter watch the trend line.
1.3 A Hard Credential — and an Irreplaceable One
In inference-deployment interviews, "talk through a systems paper" appears almost without fail. What interviewers ask is usually not the abstract, but:
- What bottleneck does this paper solve? Why is it the bottleneck?
- What is its core contribution? Why did previous methods not work?
- What do the ablations prove? Which parts are genuine algorithmic innovation, and which are just kernel tuning?
- Transplanted to my hardware (A100, not H100) and my model (13B, not 70B), where would this paper's method break?
Only those who have closely read the original can answer these. What the interviewer wants is not "you have memorized PagedAttention" but "you possess the ability to read systems papers" — because that is precisely the core skill of daily deployment optimization.
2. Where to Start: Three Routes by Goal
There is no single standard for how deeply to read a paper; pick a route by your goal. All three routes are expanded in Reading Paths; here is a quick overview first:
text
Goal: two-hour onboarding
├─ Read only FlashAttention v1 (abstract + Section 3 algorithm + Figure 1)
├─ Add one illustrated explainer or a set of notes
└─ Output: can explain "why IO-awareness makes attention 2-4x faster"
Goal: engineering deployment
├─ Deep-read vLLM / FlashAttention / GPTQ / AWQ / EAGLE-2 in full
├─ Focus on the profiles, the ablations, and the engineering limits (Discussion)
└─ Output: can judge "can this method be ported to my business model and batch size"
Goal: research
├─ Deep-read everything + math derivation + reimplement the kernel yourself
├─ Read the references backward and compare follow-ups (FA1 -> FA2 -> FA3)
└─ Output: can propose your own operator/scheduling improvement and benchmark it| Route | Who it fits | Reading depth | Papers to read |
|---|---|---|---|
| Two-hour onboarding | Learners who just finished What Is Inference Acceleration? | Abstract + introduction + algorithm figure + conclusion | Start with the FlashAttention entry in Classic Papers in Depth |
| Engineering deployment | Engineers who want papers to solve real deployment problems | Full text + ablations + limitations | Combine Classic Papers in Depth with Frontier Advances |
| Research | Systems-track graduate students planning to publish at MLSys/SOSP | Full text + derivations + kernel reimplementation | Build a global coordinate system from the Paper Map, then go deep |
One sentence for newcomers
For your first paper, always pick a "landmark and easy to read" one, not the "newest and hottest." FlashAttention v1 (NeurIPS 2022) has stood the test of time, is cited by countless follow-ups, and is clearly written. First build the confidence that "I can read a systems paper," then worry about quantity.
3. What Is in This Section
This papers section has five parts; visit them in this order:
| Page | One-line intro | What you get |
|---|---|---|
| Reading Paths | The section's main entry: specific paper lists and reading orders for the three routes | A paper-reading roadmap arranged by goal |
| Paper Map | Maps the key inference-acceleration papers of 2014-2025 by topic and timeline | A macro coordinate system: which paper is the source, which is the milestone, which is the survey |
| Classic Papers in Depth | Close-reading of the ten papers that changed the field: background, mechanism, experiments, limits | The ability to master principles, plus interview material |
| Frontier Advances | Tracks recent breakthroughs and trends: FP8/FP4, prefill/decode disaggregation, speculative decoding, new hardware | Pace with the 2024-2026 frontier |
| Reading Discipline & FAQ | Methodology for reading papers: how to take notes, what to do when stuck, how to judge a paper's quality | A sustainable reading habit, not a one-off impulse |
The logic across the section is map first, then close-reading, then frontier: the map answers "what to read," close-reading answers "how deeply," and the frontier answers "where to go next." Consult the Glossary whenever a term is unfamiliar.
4. The Single Most Important Piece of Advice
The most important advice: take away the mechanism, not the number
What sticks most easily from a paper is often the speedup figure: "2-4x faster than PyTorch," "2-3x the throughput of vLLM," "EAGLE-3 with a 6.5x speedup on MT-Bench." But remember:
Every speedup is the product of four things — model + batch + hardware + evaluation protocol. Change the model (13B instead of 70B), the hardware (A100 instead of H100), the batch size (1 instead of 256), or the evaluation script (vLLM 0.6 instead of 0.4), and the number can collapse entirely. Reproduce a 2022 FlashAttention v1 experiment on a 2024 H100 and the conclusion might even flip.
What is truly worth taking away are three things:
- The mechanism — why does this method work? What structural bottleneck does it solve? (Example: FlashAttention moves the softmax intermediate state from HBM to SRAM, defeating the HBM bandwidth wall.)
- The failure analysis — which experiments did the authors skip, on which hardware does it not apply, what limitations do the authors themselves state? Failures are often worth more than successes.
- The boundary of applicability — under what batch, what sequence length, and what hardware generation does it hold? (Example: PagedAttention pays off most at long sequences and large batches, and barely helps a single user on short sequences.)
Benchmark numbers expire; mechanisms and boundaries do not.
Further Reading
- What Is Inference Acceleration? — the foundation of concepts and paradigms; read it for the global map before reading papers
- Reading Paths — the next step from this page: pick your first route
- Paper Map — lay the key inference-acceleration papers of a decade on a timeline
- Classic Papers in Depth — section-by-section breakdowns of the ten most important papers
- Glossary — a quick terminology lookup while reading
References
The following are all real public resources for self-study:
- arXiv preprint server — where most inference-systems papers are released as preprints; focus on
cs.LG,cs.DC,cs.OS - Papers with Code — papers + code + inference-benchmark aggregation; handy for checking SOTA and reproductions
- FlashAttention official repo — the companion code of the paper; the best crutch for reading the FlashAttention series
- vLLM official repo — the industrial-strength implementation of PagedAttention; issues and PRs hide a wealth of engineering detail
- SOSP 2023 proceedings — the conference where the vLLM paper appeared; read DistServe and other systems papers from the same year
- CUDA Programming Guide — the hardware background required for reading operator papers (FlashAttention, FlashInfer)
- Tri Dao's blog — the FlashAttention author's blog, often with "author's-view" follow-up notes