Appearance
Papers: Start Here
On the surface, a deployment engineer's job looks like "get the model running, keep the service stable." But what determines your ceiling is how well you understand the principles behind the tools. Why is vLLM fast? Where do GPTQ and AWQ actually differ? Why is quantization unkind to small models? The answers to these questions are written in papers, not READMEs. This section walks you through the classic papers of the model deployment field, one paper at a time.
Why Deployment Engineers Should Read Papers
The direct payoff of reading papers comes in three layers:
- Understanding what's under the hood: vLLM's PagedAttention borrows from operating-system virtual memory; Triton's multi-model scheduling inherits Clipper's batching ideas; every algorithm in the quantization toolbox has a paper behind its trade-offs. Once you understand the principles, you graduate from "knowing which parameters to set" to "knowing why those parameters."
- Judgment when new things appear: when a new framework or algorithm shows up, you can quickly tell which old pain point it solves and how it relates to existing options — far more efficient than chasing every new blog post.
- A shared language for communication and interviews: when an interviewer asks "why AWQ instead of GPTQ?", they are really asking whether you've turned the papers into your own judgment. See Career.
If you haven't yet built a mental map of deployment, read What is model deployment and A brief history of deployment first, then come back to the papers — papers reward readers who already have a framework and are looking for detail.
How to Read a Paper: The Three-Pass Method
Don't expect to get a paper in one read. The standard academic approach is the three-pass method (from S. Keshav's How to Read a Paper, widely adopted):
| Pass | What to do | Time | What you should be able to answer afterwards |
|---|---|---|---|
| Pass 1: Skim | Read the title, abstract, section headings, figure captions, and conclusions. Answer three questions: what problem does this paper tackle? What is the core contribution? How does it relate to approaches I already know? | 5-15 minutes | "The one-sentence contribution" (every walkthrough in this section hands you that sentence) |
| Pass 2: Figures and conclusions | Study every figure and table, the experimental setup, and the key numbers; skip the derivations. Try explaining to a colleague roughly how the method works. | 30-60 minutes | The method's high-level mechanism + the key numbers (how much speedup, how much accuracy lost) |
| Pass 3: Method details | Read section by section, rederive the equations yourself, scrutinize whether the experiments are fair, and pin down the limitations. Usually only worth doing when you plan to implement or reproduce the work. | 1-3 hours | You can retell the paper end to end and point out its pitfalls |
The deployment engineer's "pass two first" strategy
You don't need to take every paper to pass three. Use pass one to decide what's worth your time and pass two to build understanding; only invest in pass three when you plan to implement the method yourself or adapt it for production. The walkthrough pages in this section have essentially done the pass-one-plus-pass-two work for you.
How to take notes:
- One card per paper:
one-sentence contribution / motivation / method in one sentence / 3 key numbers / limitations / what it means today. - Record key numbers separately, not just conclusions — numbers like "2-4x" and "99% no degradation" are the most useful ammunition in interviews and when choosing between options.
- Connect papers into a graph: how does the paper you just read relate to ones you've already read? (Orca → vLLM → SGLang is one lineage; LLM.int8 → GPTQ → AWQ is another.) The site's paper map is exactly that graph.
Recommended Reading Order by Topic
Not sure where to start? Three recommended paths, from lightest to heaviest:
| Path | Who it's for | Papers | Reading time | Entry point |
|---|---|---|---|---|
| Beginner | New to deployment, wants to build intuition | 3 papers | Half a day to 1 day | Reading path: Beginner |
| Systematic | Already doing deployment, wants to fill in the theoretical foundations | 6 papers | 2-4 days | Reading path: Systematic |
| LLM inference track | Doing or preparing to do LLM inference | 4 papers | 1-2 days | Reading path: LLM inference |
Each path lists its prerequisites (linked to the relevant Core Concepts pages), reading time, and the questions you should be able to answer afterwards. See Reading Paths.
All Paper Pages in This Section
| Page | Contents |
|---|---|
| Paper map | Topic-grouped overview: quantization, distillation, compression, serving systems, parallelism, LLM inference |
| Quantization classics: LLM.int8 / GPTQ / AWQ | Three milestone papers of LLM quantization, plus SmoothQuant |
| Distillation classics: KD and the distillation family | Hinton KD, DistilBERT, TinyBERT, MiniLM |
| PagedAttention: the vLLM system paper | Must-read from SOSP 2023: KV cache paging + continuous batching |
| Inference serving systems: Clipper / Orca / Nexus and more | The serving-systems lineage from TensorFlow Serving to Orca |
| Parallel and distributed inference | Megatron-LM, GPipe, DeepSpeed ZeRO |
About the non-walkthrough papers in the paper map
This section contains full walkthroughs of 16 core papers. The paper map also lists related papers on pruning, low-rank methods, speculative decoding, and more that are annotated but not walked through; they exist to help you build a complete picture of the field, and you can dig into any of them via the map when you need to.
Reading Discipline and FAQ
Reading discipline (for procrastinators):
- Read the abstract first, then decide whether to read the full text — for most papers you can tell within 5 minutes that they're irrelevant to you.
- Read one paper at a time, and write the paper card immediately after finishing — don't let them pile up.
- Read with questions: before you start, write down what you want to learn from the paper. It's twice as effective as reading aimlessly.
- If you can retell it, you own it: after reading, try explaining the paper to a colleague in three sentences. Whatever you can't explain is what you didn't understand.
FAQ:
- What if I can't understand a paper? Fall back through the three passes: failing to grasp the method details (pass three) is normal — passes one and two never require you to follow the equations. For the vast majority of deployment scenarios, pass two's "figures + numbers" is enough. If you're truly stuck, check the glossary and the corresponding Core Concepts page to fill in the concepts, then come back to the paper.
- How much math do I need? The papers covered here don't require graduate-level math: the quantization papers need some linear algebra intuition (matrix multiplication; treat the Hessian's "second-order information" as a black box); distillation needs cross-entropy; parallelism needs intuition for how matrices get split. Formulas are explained in words rather than rigorously derived.
- Papers vs. blog posts — which first? Blogs get you up and running with a tool quickly; papers tell you the "why" and where the boundaries are. They complement each other: use the site's case study to get the tool running, then come back to the papers to understand the trade-offs.
- Should I read the appendix and the open-source code? Not for passes one and two. It's only worth reading the code when you plan to reproduce the algorithm (say, implement a quantization method yourself) or evaluate whether a paper can be dropped straight into production.
Further Reading
- Reading Paths: how to choose between the three — Beginner / Systematic / LLM inference
- Paper map: 30+ papers, one sentence each — see the whole field at a glance
- A brief history of deployment — timeline context for the papers and their tools
- Glossary — quick reference when a paper throws an unfamiliar term at you
- Core Concepts: inference — the conceptual foundation to build before reading papers