Skip to content

Paper Map

At a glance A paper map of the model deployment field — quantization, distillation, compression, inference serving systems, parallelism and distribution, LLM inference — with representative papers grouped by topic, each with a one-sentence contribution and links to the site's walkthrough pages.

Paper Map: A Panorama of the Model Deployment Field ​

This map groups the representative papers of model deployment into 6 topics, gives each a one-sentence contribution, and notes where the site covers it in depth. It serves two purposes: before reading, use it to plan your route (what to read next, how a paper relates to what you've already read); after reading, use it to check the big picture (did I miss an entire direction?).

How to read this map

The "Site coverage" column links to a walkthrough page when one exists; entries marked "not covered" (pruning, low-rank, speculative decoding, etc.) are explained on the corresponding concept page — dig deeper there when you need to.

1. Quantization ​

Core problem: compress weights (and activations) from FP16/FP32 down to INT8/INT4 — halve memory, speed up inference, without wrecking accuracy.

PaperYear / VenueOne-sentence contributionSite coverage
LLM.int8() (Dettmers et al.)2022 / NeurIPS 2022Identified activation outliers as the cause of INT8 breakdown; a mixed-precision decomposition ("outlier features in FP16, the rest in INT8") lets a 175B model run INT8 inference on a single GPU with zero performance degradationQuantization classics
GPTQ (Frantar et al.)2023 / ICLR 2023Hessian-based layer-wise weight quantization: quantizes a 175B model to 3/4-bit at near-FP16 accuracy in about 4 GPU-hours on a single GPUQuantization classics
AWQ (Lin et al.)2023 / MLSys 2024 (Best Paper)Protects the ~1% of weights that matter, selected by activation magnitude, via equivalent scaling; hardware-friendly 4-bit quantization that generalizes wellQuantization classics
SmoothQuant (Xiao et al.)2023 / ICML 2023Shifts the quantization difficulty from activations to weights (a mathematically equivalent transform), making W8A8 viable for LLMs, with a 1.56x speedupQuantization classics

2. Knowledge Distillation ​

Core problem: use a large model's (teacher's) soft outputs to teach a small model (student) — trading training cost for inference cost.

PaperYear / VenueOne-sentence contributionSite coverage
Distilling the Knowledge in a Neural Network (Hinton et al.)2015 / NIPS 2014 DL WorkshopIntroduced distillation: soft labels plus temperature T make knowledge transferable, and the student converges better than one trained directlyDistillation classics
DistilBERT (Sanh et al.)2019 / NeurIPS 2019 EMC² WorkshopDistillation during pretraining: 40% of the parameters, 97% of the performance, 60% fasterDistillation classics
TinyBERT (Jiao et al.)2020 / EMNLP 2020 FindingsTwo-stage distillation plus attention-matrix distillation: a 4-layer model with 96.8% of the performance, 9.4x fasterDistillation classics
MiniLM (Wang et al.)2020 / EMNLP 2020 FindingsDistills the internal structure of self-attention (QK^T and V^T V): 50% of the parameters retain 99%+ accuracyDistillation classics

3. Model Compression: Pruning and Low-Rank ​

Core problem: the two compression routes besides quantization — removing unimportant weights (pruning) and replacing large matrices with low-rank approximations (low-rank factorization).

PaperYear / VenueOne-sentence contributionSite coverage
Learning both Weights and Connections for Efficient Neural Networks (Han et al.)2015 / NIPS 2015Classic pruning: train → prune small weights → retrain; compresses AlexNet/VGG 9-13x with no accuracy lossNot covered; see Model compression
Exploiting Linear Structure Within Convolutional Networks (Denton et al.)2014 / NIPS 2014Decomposes convolutional kernels into low-rank approximations with SVD: 2-3x speedup with little accuracy loss — the origin of low-rank factorizationNot covered; see Model compression
LoRA: Low-Rank Adaptation (Hu et al.)2021 / ICLR 2022Low-rank ideas enter the LLM era: freeze the original weights and train only low-rank deltas, slashing fine-tuning memory and storageNot covered; see LLM inference optimization

4. Inference Serving Systems ​

Core problem: turning a model into an online service that is high-throughput, low-latency, and manageable — how the system schedules requests and resources.

PaperYear / VenueOne-sentence contributionSite coverage
TensorFlow-Serving (Olston et al.)2016 system / 2017 paperThe first production-grade model serving framework: servable versioning, hot swapping, dynamic batching, gRPCServing systems
Clipper (Crankshaw et al.)2017 / NSDI 2017A general-purpose low-latency prediction serving layer: caching + latency-aware batching + adaptive model selectionServing systems
Nexus (Shen et al.)2019 / SOSP 2019A GPU-cluster inference engine: schedules DNNs as fragments and lets multiple applications share GPUs; 1.8-12.7x higher throughput than SOTA under latency constraintsServing systems
Orca (Yu et al.)2022 / OSDI 2022Iteration-level scheduling: swaps batches per iteration instead of per request; 36.9x higher GPT-3 175B throughput than FasterTransformer; the precursor of continuous batchingServing systems
Efficient Memory Management with PagedAttention (Kwon et al.)2023 / SOSP 2023Brings OS-style virtual memory paging to KV cache management: near-zero waste + prefix sharing, 2-4x throughputvLLM paper

5. Parallelism and Distribution ​

Core problem: the model doesn't fit on a single GPU — now what? Split it along four dimensions: data, layers, matrices, and state.

PaperYear / VenueOne-sentence contributionSite coverage
Megatron-LM (Shoeybi et al.)2019 / arXivTensor parallelism: splits layer-internal matrices by rows/columns across GPUs; 15.1 PFLOPS on 512 GPUs for an 8.3B model, 76% scaling efficiencyParallel and distributed inference
GPipe (Huang et al.)2018 / arXivPipeline parallelism: splits layers into stages across GPUs + micro-batch pipelining; near-linear speedup for a 128-layer, 6B-parameter TransformerParallel and distributed inference
DeepSpeed ZeRO (Rajbhandari et al.)2019 / arXivEliminates redundancy in three stages: optimizer states → gradients → parameters; superlinear scaling to 100B+ parameters on 400 GPUsParallel and distributed inference

6. LLM Inference ​

Core problem: the memory and scheduling challenges of autoregressive generation — the KV cache, continuous batching, and faster generation.

PaperYear / VenueOne-sentence contributionSite coverage
vLLM / PagedAttention (Kwon et al.)2023 / SOSP 2023Paged KV cache management + continuous batching; 2-4x throughput; the watershed for LLM inference systemsvLLM paper
Fast Inference from Transformers via Speculative Decoding (Leviathan et al.)2023 / ICML 2023A small model drafts, the large model verifies: 2-3x faster LLM generation with no changes to the modelNot covered; see LLM inference optimization
Accelerating LLM Decoding with Speculative Sampling (Chen et al.)2023 / arXivThe parallel variant: verifies multiple draft tokens at once, mathematically guaranteeing the output distribution is unchangedNot covered; see LLM inference optimization

The same paper can appear under multiple topics

The map is grouped by topic, so papers can appear in more than one group (e.g., PagedAttention is a landmark for both "inference serving systems" and "LLM inference"). This is intentional — the same paper teaches you different things depending on the angle you take.

Lay the tables above out by year and the field's throughlines become clear:

text
2014-2016      2017-2019            2020-2022              2022-present
┌────────────┐ ┌───────────────────┐ ┌────────────────────┐  ┌───────────────────────┐
│ Early      │→│ Compression goes  │→│ Parallelism & LLM  │→ │ LLM inference systems │
│ compression│ │ practical         │ │ training           │  │ Orca→vLLM             │
│ Pruning/   │ │ Serving systems   │ │ Megatron/GPipe/    │  │ LLM.int8/GPTQ/        │
│ distill/   │ │ TFServing/Clipper │ │ ZeRO               │  │ AWQ/SmoothQuant       │
│ low-rank/  │ │ DistilBERT        │ │ GPTQ/AWQ brewing   │  │ Spec. dec./prefix     │
│ SVD        │ │ Quant. as engineer│ │                    │  │ caching               │
└────────────┘ └───────────────────┘ └────────────────────┘  └───────────────────────┘
   Compression    From "algorithms"     The scaling war      The efficiency war
   begins         to systems            on training          on inference

Four trends:

  1. From "making models smaller" to "making systems smarter": from 2015 to 2019 the stars were compression algorithms (pruning, distillation, quantization); after 2022 they became schedulers and memory managers (Orca, vLLM). When the model itself can't get any smaller, the battle is decided at the system level.
  2. Training-side techniques are migrating wholesale to inference: tensor parallelism (Megatron) and memory partitioning (ZeRO) began as training techniques and are now standard configuration for LLM inference; pipeline parallelism is being absorbed by inference engines too.
  3. LLM quantization covered an entire era in two years: in 2022 LLM.int8 merely avoided breaking; in 2023 GPTQ/AWQ pushed 4-bit to near-lossless and SmoothQuant made W8A8 production-viable — quantization went from "is it possible?" to "which one do you pick?".
  4. The KV cache has become new infrastructure: after PagedAttention, prefix caching, chunked prefill, and KV compression (the GQA line of work, among others) became hot spots for new papers — exactly the battleground deployment engineers should keep tracking over the next few years. See LLM inference optimization.

Further Reading ​