Appearance
Paper Map: A Panorama of the Model Deployment Field
This map groups the representative papers of model deployment into 6 topics, gives each a one-sentence contribution, and notes where the site covers it in depth. It serves two purposes: before reading, use it to plan your route (what to read next, how a paper relates to what you've already read); after reading, use it to check the big picture (did I miss an entire direction?).
How to read this map
The "Site coverage" column links to a walkthrough page when one exists; entries marked "not covered" (pruning, low-rank, speculative decoding, etc.) are explained on the corresponding concept page — dig deeper there when you need to.
1. Quantization
Core problem: compress weights (and activations) from FP16/FP32 down to INT8/INT4 — halve memory, speed up inference, without wrecking accuracy.
| Paper | Year / Venue | One-sentence contribution | Site coverage |
|---|---|---|---|
| LLM.int8() (Dettmers et al.) | 2022 / NeurIPS 2022 | Identified activation outliers as the cause of INT8 breakdown; a mixed-precision decomposition ("outlier features in FP16, the rest in INT8") lets a 175B model run INT8 inference on a single GPU with zero performance degradation | Quantization classics |
| GPTQ (Frantar et al.) | 2023 / ICLR 2023 | Hessian-based layer-wise weight quantization: quantizes a 175B model to 3/4-bit at near-FP16 accuracy in about 4 GPU-hours on a single GPU | Quantization classics |
| AWQ (Lin et al.) | 2023 / MLSys 2024 (Best Paper) | Protects the ~1% of weights that matter, selected by activation magnitude, via equivalent scaling; hardware-friendly 4-bit quantization that generalizes well | Quantization classics |
| SmoothQuant (Xiao et al.) | 2023 / ICML 2023 | Shifts the quantization difficulty from activations to weights (a mathematically equivalent transform), making W8A8 viable for LLMs, with a 1.56x speedup | Quantization classics |
2. Knowledge Distillation
Core problem: use a large model's (teacher's) soft outputs to teach a small model (student) — trading training cost for inference cost.
| Paper | Year / Venue | One-sentence contribution | Site coverage |
|---|---|---|---|
| Distilling the Knowledge in a Neural Network (Hinton et al.) | 2015 / NIPS 2014 DL Workshop | Introduced distillation: soft labels plus temperature T make knowledge transferable, and the student converges better than one trained directly | Distillation classics |
| DistilBERT (Sanh et al.) | 2019 / NeurIPS 2019 EMC² Workshop | Distillation during pretraining: 40% of the parameters, 97% of the performance, 60% faster | Distillation classics |
| TinyBERT (Jiao et al.) | 2020 / EMNLP 2020 Findings | Two-stage distillation plus attention-matrix distillation: a 4-layer model with 96.8% of the performance, 9.4x faster | Distillation classics |
| MiniLM (Wang et al.) | 2020 / EMNLP 2020 Findings | Distills the internal structure of self-attention (QK^T and V^T V): 50% of the parameters retain 99%+ accuracy | Distillation classics |
3. Model Compression: Pruning and Low-Rank
Core problem: the two compression routes besides quantization — removing unimportant weights (pruning) and replacing large matrices with low-rank approximations (low-rank factorization).
| Paper | Year / Venue | One-sentence contribution | Site coverage |
|---|---|---|---|
| Learning both Weights and Connections for Efficient Neural Networks (Han et al.) | 2015 / NIPS 2015 | Classic pruning: train → prune small weights → retrain; compresses AlexNet/VGG 9-13x with no accuracy loss | Not covered; see Model compression |
| Exploiting Linear Structure Within Convolutional Networks (Denton et al.) | 2014 / NIPS 2014 | Decomposes convolutional kernels into low-rank approximations with SVD: 2-3x speedup with little accuracy loss — the origin of low-rank factorization | Not covered; see Model compression |
| LoRA: Low-Rank Adaptation (Hu et al.) | 2021 / ICLR 2022 | Low-rank ideas enter the LLM era: freeze the original weights and train only low-rank deltas, slashing fine-tuning memory and storage | Not covered; see LLM inference optimization |
4. Inference Serving Systems
Core problem: turning a model into an online service that is high-throughput, low-latency, and manageable — how the system schedules requests and resources.
| Paper | Year / Venue | One-sentence contribution | Site coverage |
|---|---|---|---|
| TensorFlow-Serving (Olston et al.) | 2016 system / 2017 paper | The first production-grade model serving framework: servable versioning, hot swapping, dynamic batching, gRPC | Serving systems |
| Clipper (Crankshaw et al.) | 2017 / NSDI 2017 | A general-purpose low-latency prediction serving layer: caching + latency-aware batching + adaptive model selection | Serving systems |
| Nexus (Shen et al.) | 2019 / SOSP 2019 | A GPU-cluster inference engine: schedules DNNs as fragments and lets multiple applications share GPUs; 1.8-12.7x higher throughput than SOTA under latency constraints | Serving systems |
| Orca (Yu et al.) | 2022 / OSDI 2022 | Iteration-level scheduling: swaps batches per iteration instead of per request; 36.9x higher GPT-3 175B throughput than FasterTransformer; the precursor of continuous batching | Serving systems |
| Efficient Memory Management with PagedAttention (Kwon et al.) | 2023 / SOSP 2023 | Brings OS-style virtual memory paging to KV cache management: near-zero waste + prefix sharing, 2-4x throughput | vLLM paper |
5. Parallelism and Distribution
Core problem: the model doesn't fit on a single GPU — now what? Split it along four dimensions: data, layers, matrices, and state.
| Paper | Year / Venue | One-sentence contribution | Site coverage |
|---|---|---|---|
| Megatron-LM (Shoeybi et al.) | 2019 / arXiv | Tensor parallelism: splits layer-internal matrices by rows/columns across GPUs; 15.1 PFLOPS on 512 GPUs for an 8.3B model, 76% scaling efficiency | Parallel and distributed inference |
| GPipe (Huang et al.) | 2018 / arXiv | Pipeline parallelism: splits layers into stages across GPUs + micro-batch pipelining; near-linear speedup for a 128-layer, 6B-parameter Transformer | Parallel and distributed inference |
| DeepSpeed ZeRO (Rajbhandari et al.) | 2019 / arXiv | Eliminates redundancy in three stages: optimizer states → gradients → parameters; superlinear scaling to 100B+ parameters on 400 GPUs | Parallel and distributed inference |
6. LLM Inference
Core problem: the memory and scheduling challenges of autoregressive generation — the KV cache, continuous batching, and faster generation.
| Paper | Year / Venue | One-sentence contribution | Site coverage |
|---|---|---|---|
| vLLM / PagedAttention (Kwon et al.) | 2023 / SOSP 2023 | Paged KV cache management + continuous batching; 2-4x throughput; the watershed for LLM inference systems | vLLM paper |
| Fast Inference from Transformers via Speculative Decoding (Leviathan et al.) | 2023 / ICML 2023 | A small model drafts, the large model verifies: 2-3x faster LLM generation with no changes to the model | Not covered; see LLM inference optimization |
| Accelerating LLM Decoding with Speculative Sampling (Chen et al.) | 2023 / arXiv | The parallel variant: verifies multiple draft tokens at once, mathematically guaranteeing the output distribution is unchanged | Not covered; see LLM inference optimization |
The same paper can appear under multiple topics
The map is grouped by topic, so papers can appear in more than one group (e.g., PagedAttention is a landmark for both "inference serving systems" and "LLM inference"). This is intentional — the same paper teaches you different things depending on the angle you take.
Trends on a Timeline
Lay the tables above out by year and the field's throughlines become clear:
text
2014-2016 2017-2019 2020-2022 2022-present
┌────────────┐ ┌───────────────────┐ ┌────────────────────┐ ┌───────────────────────┐
│ Early │→│ Compression goes │→│ Parallelism & LLM │→ │ LLM inference systems │
│ compression│ │ practical │ │ training │ │ Orca→vLLM │
│ Pruning/ │ │ Serving systems │ │ Megatron/GPipe/ │ │ LLM.int8/GPTQ/ │
│ distill/ │ │ TFServing/Clipper │ │ ZeRO │ │ AWQ/SmoothQuant │
│ low-rank/ │ │ DistilBERT │ │ GPTQ/AWQ brewing │ │ Spec. dec./prefix │
│ SVD │ │ Quant. as engineer│ │ │ │ caching │
└────────────┘ └───────────────────┘ └────────────────────┘ └───────────────────────┘
Compression From "algorithms" The scaling war The efficiency war
begins to systems on training on inferenceFour trends:
- From "making models smaller" to "making systems smarter": from 2015 to 2019 the stars were compression algorithms (pruning, distillation, quantization); after 2022 they became schedulers and memory managers (Orca, vLLM). When the model itself can't get any smaller, the battle is decided at the system level.
- Training-side techniques are migrating wholesale to inference: tensor parallelism (Megatron) and memory partitioning (ZeRO) began as training techniques and are now standard configuration for LLM inference; pipeline parallelism is being absorbed by inference engines too.
- LLM quantization covered an entire era in two years: in 2022 LLM.int8 merely avoided breaking; in 2023 GPTQ/AWQ pushed 4-bit to near-lossless and SmoothQuant made W8A8 production-viable — quantization went from "is it possible?" to "which one do you pick?".
- The KV cache has become new infrastructure: after PagedAttention, prefix caching, chunked prefill, and KV compression (the GQA line of work, among others) became hot spots for new papers — exactly the battleground deployment engineers should keep tracking over the next few years. See LLM inference optimization.
Further Reading
- Papers: Start Here — the three-pass method and note template
- Reading Paths — three topic-organized reading routes
- Quantization classics: LLM.int8 / GPTQ / AWQ
- Distillation classics: KD and the distillation family
- PagedAttention: the vLLM system paper
- Inference serving systems: Clipper / Orca / Nexus and more
- Parallel and distributed inference
- A brief history of deployment — the map's timeline view alongside how the tools evolved