Appearance
Reading Paths: Three Paper Routes
This section currently walks through 16 papers, and reading them all in order isn't realistic. This page splits them into three paths: Beginner (build intuition), Systematic (fill in the foundations), and LLM Inference (for serving large models). Pick one, finish it, then decide whether to continue.
The Three Paths at a Glance
| Path | Goal | Papers | Suggested time | Who it's for |
|---|---|---|---|---|
| Beginner | Build a feel for "deployment + papers"; be able to state each paper's one-sentence contribution | 3 papers | Half a day to 1 day | Newcomers to deployment, readers who want concepts fast |
| Systematic | Systematically cover the theory behind quantization, distillation, and parallelism; be able to make technology choices | 6 papers | 2-4 days | Engineers with deployment experience who want the principles |
| LLM inference track | Master the four main threads of LLM inference: scheduling, memory, quantization, parallelism | 4 papers | 1-2 days | Anyone doing or preparing to do LLM inference |
Paths can overlap
The split is organized by topic, not by strictly increasing difficulty. TensorFlow Serving in the beginner path isn't actually shallow — it just carries a light conceptual load. Once you finish the beginner path, pick either the systematic or the LLM-inference route based on interest; you don't need to do both.
Beginner (3 papers)
Order: Hinton KD → LLM.int8() → TensorFlow Serving
Why this order: first understand how models get smaller (distillation → quantization), then how a model becomes a service (serving systems). Each step's concepts are the foundation for the next.
1. Distilling the Knowledge in a Neural Network (Hinton 2015)
- One-sentence contribution: the soft labels a large model outputs carry more information than hard labels; training a small model on them lets it converge better.
- Prerequisites: Model compression (what distillation is), Inference basics (basic training/inference concepts).
- Reading time: 30-45 minutes (the distillation walkthrough explains the formulas in plain language).
- You should be able to answer: what extra information do soft labels give a model beyond hard labels? What do higher and lower values of the temperature T each mean?
2. LLM.int8() (Dettmers 2022)
- One-sentence contribution: identified outliers in LLM activations as the cause of INT8 breakdown, and used mixed-precision decomposition to run a 175B model on a single GPU in INT8 with zero performance degradation.
- Prerequisites: Quantization (how INT8 quantization works, where accuracy loss comes from), GPUs and hardware (why memory is the bottleneck).
- Reading time: 45-60 minutes (quantization walkthrough).
- You should be able to answer: why does quantization work fine on small models but break large ones? Which part does mixed-precision decomposition keep in FP16, and what fraction of the total is it?
3. TensorFlow Serving (Olston 2017)
- One-sentence contribution: the first production-grade model serving framework, codifying "model versioning + hot swapping + dynamic batching + gRPC" as a paradigm.
- Prerequisites: Serving and inference APIs, Deployment architecture patterns.
- Reading time: 30-45 minutes (serving systems walkthrough).
- You should be able to answer: what problem does servable version management solve? Why does dynamic batching raise throughput, and what does it cost?
How to check you've finished the beginner path
After the three papers, you pass if you can explain to a colleague which route distillation and quantization each take to shrink a model, and why virtually every inference service today batches requests. Then make the concepts hands-on with Deploy a model from scratch.
Systematic Track (6 papers)
Order: DistilBERT → GPTQ → AWQ → Megatron-LM → GPipe → DeepSpeed ZeRO
Thematic thread: first finish both branches of model compression (practical distillation + the two mainstream quantization methods), then switch to the parallelism trio for "the model doesn't fit" (tensor parallelism → pipeline parallelism → memory partitioning).
| Paper | Prerequisites | Time | What you should be able to answer |
|---|---|---|---|
| DistilBERT (Sanh 2019) | Model compression + having read KD | 45-60 minutes | How does distilling during pretraining differ from task-level distillation? What exactly do 40% of the parameters, 97% of the performance, and 60% faster refer to? |
| GPTQ (Frantar 2023) | Quantization + having read LLM.int8 | 45-60 minutes | Why is Hessian-based layer-wise quantization more accurate than per-element scaling? How close do 3-bit/4-bit get to FP16? |
| AWQ (Lin 2023) | Quantization + having read GPTQ | 45-60 minutes | How do AWQ and GPTQ differ in what they're for? Why isn't "protecting the 1% of important weights" the same thing as mixed precision? |
| Megatron-LM (Shoeybi 2020) | GPUs and hardware + LLM inference | 60-90 minutes | How does tensor parallelism split the matrices, and how many communications happen per forward pass? Why does it suit inference, not just training? |
| GPipe (Huang 2019) | GPUs and hardware + Deployment patterns | 45-60 minutes | How do micro-batches eliminate pipeline idle time? How does bubble overhead relate to the number of layers and micro-batches? |
| DeepSpeed ZeRO (Rajbhandari 2020) | GPUs and hardware + LLM inference | 60-90 minutes | Which redundancy does each of ZeRO's three stages eliminate? Why is Stage 3 (parameter partitioning) the most common choice for inference? |
How the systematic path fits together
After six papers, a decision logic starts to take shape: model too large → quantize (GPTQ/AWQ) + parallelize (Megatron/GPipe/ZeRO); latency-sensitive → distill into a small model (DistilBERT). These judgments map directly onto the engineering conclusions in the Quantization and Model compression pages.
LLM Inference Track (4 papers)
Order: Orca → vLLM (PagedAttention) → SmoothQuant → DeepSpeed ZeRO
Main thread: nail scheduling first (Orca predates vLLM's continuous batching), then read the vLLM system paper itself, then fill in two supporting threads — quantization (SmoothQuant makes W8A8 viable) and memory (ZeRO partitioning).
| Paper | Prerequisites | Time | What you should be able to answer |
|---|---|---|---|
| Orca (Yu 2022) | LLM inference + Serving | 60-90 minutes | How does iteration-level scheduling differ from request-level scheduling? Why do LLMs need the former? |
| vLLM / PagedAttention (Kwon 2023) | LLM inference + having read Orca | 60-90 minutes | Why does the KV cache waste 60-80% of its memory? How does the block table eliminate fragmentation and support prefix sharing? |
| SmoothQuant (Xiao 2023) | Quantization | 45-60 minutes | Why are activations harder to quantize than weights? When you shift the quantization difficulty from activations to weights, why is it "mathematically equivalent and engineering-feasible"? |
| DeepSpeed ZeRO (Rajbhandari 2020) | GPUs and hardware + LLM inference | 60-90 minutes | What does ZeRO mainly save on the inference side? Is it the same layer of problem as vLLM's KV cache optimization? (No: one is about weights, the other about the KV cache.) |
The highlight of this track
The vLLM paper (SOSP 2023) is the only walkthrough in this section marked recommended — it's the template for "operating-system ideas rescuing LLM inference," and the case study maps one-to-one onto the paper. Read the vLLM paper before picking up vLLM and you'll notice that every parameter in the docs has a paper behind it.
What About Papers Outside the Paths
The three paths cover most of the 16 walkthroughs; read the remaining ones as needed:
- TinyBERT and MiniLM: read when doing NLP model compression or actually distilling down to a small model. See the distillation walkthrough.
- Clipper and Nexus: read when you care about classic inference serving or GPU multi-tenant scheduling. See the serving systems walkthrough.
- The paper map also lists papers on pruning, low-rank methods, speculative decoding, and more that are annotated but not walked through — on-demand extensions.
After You Finish the Paths
Revisit the Papers home page to review the three-pass note template, then put the papers' conclusions into practice:
- Deploy a model from scratch — connect "shrink the model + serve it" into one hands-on route
- vLLM for LLM inference — see real configurations alongside the vLLM paper
- Choosing frameworks and platforms — make choices using the trade-off logic from the papers
- Common pitfalls and anti-patterns — every "limitation" in a paper is one you'll hit in production
Further Reading
- Papers: Start Here — the three-pass method, note template, and FAQ
- Paper map — all topic groups and one-sentence contributions
- Quantization classics: LLM.int8 / GPTQ / AWQ
- Distillation classics: KD and the distillation family
- PagedAttention: the vLLM system paper
- Parallel and distributed inference