Skip to content

Papers: Start Here

At a glance Why should deployment engineers read papers? This page explains how to use the Papers section: the three-pass reading method, recommended reading order by topic, and entry points to every walkthrough.

Papers: Start Here ​

On the surface, a deployment engineer's job looks like "get the model running, keep the service stable." But what determines your ceiling is how well you understand the principles behind the tools. Why is vLLM fast? Where do GPTQ and AWQ actually differ? Why is quantization unkind to small models? The answers to these questions are written in papers, not READMEs. This section walks you through the classic papers of the model deployment field, one paper at a time.

Why Deployment Engineers Should Read Papers ​

The direct payoff of reading papers comes in three layers:

  1. Understanding what's under the hood: vLLM's PagedAttention borrows from operating-system virtual memory; Triton's multi-model scheduling inherits Clipper's batching ideas; every algorithm in the quantization toolbox has a paper behind its trade-offs. Once you understand the principles, you graduate from "knowing which parameters to set" to "knowing why those parameters."
  2. Judgment when new things appear: when a new framework or algorithm shows up, you can quickly tell which old pain point it solves and how it relates to existing options — far more efficient than chasing every new blog post.
  3. A shared language for communication and interviews: when an interviewer asks "why AWQ instead of GPTQ?", they are really asking whether you've turned the papers into your own judgment. See Career.

If you haven't yet built a mental map of deployment, read What is model deployment and A brief history of deployment first, then come back to the papers — papers reward readers who already have a framework and are looking for detail.

How to Read a Paper: The Three-Pass Method ​

Don't expect to get a paper in one read. The standard academic approach is the three-pass method (from S. Keshav's How to Read a Paper, widely adopted):

PassWhat to doTimeWhat you should be able to answer afterwards
Pass 1: SkimRead the title, abstract, section headings, figure captions, and conclusions. Answer three questions: what problem does this paper tackle? What is the core contribution? How does it relate to approaches I already know?5-15 minutes"The one-sentence contribution" (every walkthrough in this section hands you that sentence)
Pass 2: Figures and conclusionsStudy every figure and table, the experimental setup, and the key numbers; skip the derivations. Try explaining to a colleague roughly how the method works.30-60 minutesThe method's high-level mechanism + the key numbers (how much speedup, how much accuracy lost)
Pass 3: Method detailsRead section by section, rederive the equations yourself, scrutinize whether the experiments are fair, and pin down the limitations. Usually only worth doing when you plan to implement or reproduce the work.1-3 hoursYou can retell the paper end to end and point out its pitfalls

The deployment engineer's "pass two first" strategy

You don't need to take every paper to pass three. Use pass one to decide what's worth your time and pass two to build understanding; only invest in pass three when you plan to implement the method yourself or adapt it for production. The walkthrough pages in this section have essentially done the pass-one-plus-pass-two work for you.

How to take notes:

  • One card per paper: one-sentence contribution / motivation / method in one sentence / 3 key numbers / limitations / what it means today.
  • Record key numbers separately, not just conclusions — numbers like "2-4x" and "99% no degradation" are the most useful ammunition in interviews and when choosing between options.
  • Connect papers into a graph: how does the paper you just read relate to ones you've already read? (Orca → vLLM → SGLang is one lineage; LLM.int8 → GPTQ → AWQ is another.) The site's paper map is exactly that graph.

Not sure where to start? Three recommended paths, from lightest to heaviest:

PathWho it's forPapersReading timeEntry point
BeginnerNew to deployment, wants to build intuition3 papersHalf a day to 1 dayReading path: Beginner
SystematicAlready doing deployment, wants to fill in the theoretical foundations6 papers2-4 daysReading path: Systematic
LLM inference trackDoing or preparing to do LLM inference4 papers1-2 daysReading path: LLM inference

Each path lists its prerequisites (linked to the relevant Core Concepts pages), reading time, and the questions you should be able to answer afterwards. See Reading Paths.

All Paper Pages in This Section ​

PageContents
Paper mapTopic-grouped overview: quantization, distillation, compression, serving systems, parallelism, LLM inference
Quantization classics: LLM.int8 / GPTQ / AWQThree milestone papers of LLM quantization, plus SmoothQuant
Distillation classics: KD and the distillation familyHinton KD, DistilBERT, TinyBERT, MiniLM
PagedAttention: the vLLM system paperMust-read from SOSP 2023: KV cache paging + continuous batching
Inference serving systems: Clipper / Orca / Nexus and moreThe serving-systems lineage from TensorFlow Serving to Orca
Parallel and distributed inferenceMegatron-LM, GPipe, DeepSpeed ZeRO

About the non-walkthrough papers in the paper map

This section contains full walkthroughs of 16 core papers. The paper map also lists related papers on pruning, low-rank methods, speculative decoding, and more that are annotated but not walked through; they exist to help you build a complete picture of the field, and you can dig into any of them via the map when you need to.

Reading Discipline and FAQ ​

Reading discipline (for procrastinators):

  • Read the abstract first, then decide whether to read the full text — for most papers you can tell within 5 minutes that they're irrelevant to you.
  • Read one paper at a time, and write the paper card immediately after finishing — don't let them pile up.
  • Read with questions: before you start, write down what you want to learn from the paper. It's twice as effective as reading aimlessly.
  • If you can retell it, you own it: after reading, try explaining the paper to a colleague in three sentences. Whatever you can't explain is what you didn't understand.

FAQ:

  • What if I can't understand a paper? Fall back through the three passes: failing to grasp the method details (pass three) is normal — passes one and two never require you to follow the equations. For the vast majority of deployment scenarios, pass two's "figures + numbers" is enough. If you're truly stuck, check the glossary and the corresponding Core Concepts page to fill in the concepts, then come back to the paper.
  • How much math do I need? The papers covered here don't require graduate-level math: the quantization papers need some linear algebra intuition (matrix multiplication; treat the Hessian's "second-order information" as a black box); distillation needs cross-entropy; parallelism needs intuition for how matrices get split. Formulas are explained in words rather than rigorously derived.
  • Papers vs. blog posts — which first? Blogs get you up and running with a tool quickly; papers tell you the "why" and where the boundaries are. They complement each other: use the site's case study to get the tool running, then come back to the papers to understand the trade-offs.
  • Should I read the appendix and the open-source code? Not for passes one and two. It's only worth reading the code when you plan to reproduce the algorithm (say, implement a quantization method yourself) or evaluate whether a paper can be dropped straight into production.

Further Reading ​