Skip to content

Reading Paths

At a glance Too many papers, not enough time? This page lays out three reading paths — Beginner (3 papers), Systematic (6 papers), and the LLM Inference track (4 papers) — each with its prerequisites and reading order.

Reading Paths: Three Paper Routes ​

This section currently walks through 16 papers, and reading them all in order isn't realistic. This page splits them into three paths: Beginner (build intuition), Systematic (fill in the foundations), and LLM Inference (for serving large models). Pick one, finish it, then decide whether to continue.

The Three Paths at a Glance ​

PathGoalPapersSuggested timeWho it's for
BeginnerBuild a feel for "deployment + papers"; be able to state each paper's one-sentence contribution3 papersHalf a day to 1 dayNewcomers to deployment, readers who want concepts fast
SystematicSystematically cover the theory behind quantization, distillation, and parallelism; be able to make technology choices6 papers2-4 daysEngineers with deployment experience who want the principles
LLM inference trackMaster the four main threads of LLM inference: scheduling, memory, quantization, parallelism4 papers1-2 daysAnyone doing or preparing to do LLM inference

Paths can overlap

The split is organized by topic, not by strictly increasing difficulty. TensorFlow Serving in the beginner path isn't actually shallow — it just carries a light conceptual load. Once you finish the beginner path, pick either the systematic or the LLM-inference route based on interest; you don't need to do both.

Beginner (3 papers) ​

Order: Hinton KD → LLM.int8() → TensorFlow Serving

Why this order: first understand how models get smaller (distillation → quantization), then how a model becomes a service (serving systems). Each step's concepts are the foundation for the next.

1. Distilling the Knowledge in a Neural Network (Hinton 2015) ​

  • One-sentence contribution: the soft labels a large model outputs carry more information than hard labels; training a small model on them lets it converge better.
  • Prerequisites: Model compression (what distillation is), Inference basics (basic training/inference concepts).
  • Reading time: 30-45 minutes (the distillation walkthrough explains the formulas in plain language).
  • You should be able to answer: what extra information do soft labels give a model beyond hard labels? What do higher and lower values of the temperature T each mean?

2. LLM.int8() (Dettmers 2022) ​

  • One-sentence contribution: identified outliers in LLM activations as the cause of INT8 breakdown, and used mixed-precision decomposition to run a 175B model on a single GPU in INT8 with zero performance degradation.
  • Prerequisites: Quantization (how INT8 quantization works, where accuracy loss comes from), GPUs and hardware (why memory is the bottleneck).
  • Reading time: 45-60 minutes (quantization walkthrough).
  • You should be able to answer: why does quantization work fine on small models but break large ones? Which part does mixed-precision decomposition keep in FP16, and what fraction of the total is it?

3. TensorFlow Serving (Olston 2017) ​

  • One-sentence contribution: the first production-grade model serving framework, codifying "model versioning + hot swapping + dynamic batching + gRPC" as a paradigm.
  • Prerequisites: Serving and inference APIs, Deployment architecture patterns.
  • Reading time: 30-45 minutes (serving systems walkthrough).
  • You should be able to answer: what problem does servable version management solve? Why does dynamic batching raise throughput, and what does it cost?

How to check you've finished the beginner path

After the three papers, you pass if you can explain to a colleague which route distillation and quantization each take to shrink a model, and why virtually every inference service today batches requests. Then make the concepts hands-on with Deploy a model from scratch.

Systematic Track (6 papers) ​

Order: DistilBERT → GPTQ → AWQ → Megatron-LM → GPipe → DeepSpeed ZeRO

Thematic thread: first finish both branches of model compression (practical distillation + the two mainstream quantization methods), then switch to the parallelism trio for "the model doesn't fit" (tensor parallelism → pipeline parallelism → memory partitioning).

PaperPrerequisitesTimeWhat you should be able to answer
DistilBERT (Sanh 2019)Model compression + having read KD45-60 minutesHow does distilling during pretraining differ from task-level distillation? What exactly do 40% of the parameters, 97% of the performance, and 60% faster refer to?
GPTQ (Frantar 2023)Quantization + having read LLM.int845-60 minutesWhy is Hessian-based layer-wise quantization more accurate than per-element scaling? How close do 3-bit/4-bit get to FP16?
AWQ (Lin 2023)Quantization + having read GPTQ45-60 minutesHow do AWQ and GPTQ differ in what they're for? Why isn't "protecting the 1% of important weights" the same thing as mixed precision?
Megatron-LM (Shoeybi 2020)GPUs and hardware + LLM inference60-90 minutesHow does tensor parallelism split the matrices, and how many communications happen per forward pass? Why does it suit inference, not just training?
GPipe (Huang 2019)GPUs and hardware + Deployment patterns45-60 minutesHow do micro-batches eliminate pipeline idle time? How does bubble overhead relate to the number of layers and micro-batches?
DeepSpeed ZeRO (Rajbhandari 2020)GPUs and hardware + LLM inference60-90 minutesWhich redundancy does each of ZeRO's three stages eliminate? Why is Stage 3 (parameter partitioning) the most common choice for inference?

How the systematic path fits together

After six papers, a decision logic starts to take shape: model too large → quantize (GPTQ/AWQ) + parallelize (Megatron/GPipe/ZeRO); latency-sensitive → distill into a small model (DistilBERT). These judgments map directly onto the engineering conclusions in the Quantization and Model compression pages.

LLM Inference Track (4 papers) ​

Order: Orca → vLLM (PagedAttention) → SmoothQuant → DeepSpeed ZeRO

Main thread: nail scheduling first (Orca predates vLLM's continuous batching), then read the vLLM system paper itself, then fill in two supporting threads — quantization (SmoothQuant makes W8A8 viable) and memory (ZeRO partitioning).

PaperPrerequisitesTimeWhat you should be able to answer
Orca (Yu 2022)LLM inference + Serving60-90 minutesHow does iteration-level scheduling differ from request-level scheduling? Why do LLMs need the former?
vLLM / PagedAttention (Kwon 2023)LLM inference + having read Orca60-90 minutesWhy does the KV cache waste 60-80% of its memory? How does the block table eliminate fragmentation and support prefix sharing?
SmoothQuant (Xiao 2023)Quantization45-60 minutesWhy are activations harder to quantize than weights? When you shift the quantization difficulty from activations to weights, why is it "mathematically equivalent and engineering-feasible"?
DeepSpeed ZeRO (Rajbhandari 2020)GPUs and hardware + LLM inference60-90 minutesWhat does ZeRO mainly save on the inference side? Is it the same layer of problem as vLLM's KV cache optimization? (No: one is about weights, the other about the KV cache.)

The highlight of this track

The vLLM paper (SOSP 2023) is the only walkthrough in this section marked recommended — it's the template for "operating-system ideas rescuing LLM inference," and the case study maps one-to-one onto the paper. Read the vLLM paper before picking up vLLM and you'll notice that every parameter in the docs has a paper behind it.

What About Papers Outside the Paths ​

The three paths cover most of the 16 walkthroughs; read the remaining ones as needed:

  • TinyBERT and MiniLM: read when doing NLP model compression or actually distilling down to a small model. See the distillation walkthrough.
  • Clipper and Nexus: read when you care about classic inference serving or GPU multi-tenant scheduling. See the serving systems walkthrough.
  • The paper map also lists papers on pruning, low-rank methods, speculative decoding, and more that are annotated but not walked through — on-demand extensions.

After You Finish the Paths ​

Revisit the Papers home page to review the three-pass note template, then put the papers' conclusions into practice:

Further Reading ​