Appearance
LLM Inference Optimization
One-line definition: LLM inference optimization is the systems engineering of combining memory management, batch scheduling, parallelism, and quantization around one defining property — autoregressive, token-by-token generation — the goal is to make every generated token fast, cheap, and able to serve more concurrent requests.
Industry insight: LLM inference rewrote the playbook of traditional inference optimization. Your intuition that "the weights dominate GPU memory" is wrong in the LLM world — as concurrency and context length grow, the KV cache can eat more memory than the weights do. "Bigger batch means better economics" breaks down too: requests have different lengths, and static batching forces the whole batch to wait for the slowest one. vLLM made its name by turning "memory management," a long-ignored problem, into its core innovation (PagedAttention), proving that systems-level engineering can deliver order-of-magnitude gains. This section gives you a complete "GPU memory ledger + optimization toolbox."
1. Three Properties of LLM Inference
- Autoregressive generation: one forward pass produces one token, and each next token depends on the previous one — generating N tokens takes N forward passes; you can't compute it all at once;
- The KV cache eats memory: every request keeps caching attention keys and values as it generates, growing linearly with sequence length;
- Decoding is bandwidth-bound: each step computes only one new token, so shipping weights through memory dominates — which is exactly why quantization and batching are so effective for LLMs.
text
LLM generation timeline:
request: "deploy the model" + generating...
Prefill: one forward pass over the full prompt → produces the first token (TTFT)
Decode : one forward pass per step → one token (TPOT, repeated for every token)2. The GPU Memory Ledger: Weights + Activations + KV Cache
Take a 7B FP16 model (weights ≈ 14GB) deployed on an 80GB A100. The memory budget breaks down like this:
text
┌────────────────────────────────────────────────────────────────┐
│ Weights ~14GB (FP16) │
│ Activations ~2GB (at batch=16) │
│ KV cache: grows with concurrency × context length │
│ per request, per token ≈ 2 × layers × heads × head_dim × 2B │
│ example: 32 layers × 128 heads × 128 dim × 2B ≈ 2MB/token │
│ 128 concurrent requests, 2K tokens context on average │
│ ≈ 128 × 2048 × 2MB = 512GB (far beyond the weights!) │
└────────────────────────────────────────────────────────────────┘The key takeaway: the KV cache is not a minor overhead — it is the dominant memory item. A formula for estimating KV cache size:
text
KV cache size = 2 × layers × num_kv_heads × head_dim × bytes (FP16=2) × total sequence lengthThat's why "how long a context and how many concurrent requests can we support" is the central memory decision in LLM deployment. Mitigations: KV cache quantization (INT8/FP4), GQA/MQA architectures (shared KV heads), and the PagedAttention-style dynamic management covered next.
3. Seven Core Optimizations
1. KV Cache Management: PagedAttention
The core idea of the vLLM paper (SOSP 2023): manage the KV cache the way an operating system manages memory — allocate fixed-size blocks (pages), link them through an index table, and eliminate the fragmentation waste of "reserving a large contiguous region of GPU memory but using only a fraction of it." The payoffs:
- Memory utilization climbs from ~60% to 90%+;
- The same page can be shared across requests (e.g., common prefixes in beam search);
- Throughput gains of up to 20×+ (measured in the paper, long-context scenarios). For a deep dive into the paper, see Papers: PagedAttention.
2. Continuous Batching
With static batching, a batch holds its resources until every request in it finishes, so short requests wait on long ones. Continuous batching switches to token-level scheduling:
text
Static batching: request A ████████
request B ████████████ ← B is short but must wait for A; the GPU idles
Continuous batching: request A ████ ✓ (finishes first, leaves first)
request B ████████
request C ████████ ← new requests slot into freed capacity immediatelyThis is the number-one lever for multiplying LLM serving throughput several times over, and vLLM, TGI, and SGLang all ship it built in.
3. Speculative Decoding
The idea: a small draft model guesses the next K tokens in one shot, then the large model verifies all of them in a single forward pass — correct guesses are pure profit; wrong ones get rolled back. When QPS isn't your constraint (you have spare compute), speedups of 2–3× are achievable, and the output matches greedy decoding exactly.
text
Draft model: guesses 4 tokens, e.g. "hello there world today"
Large model, one forward pass: verifies all 4 positions at once
→ 3 correct, 1 wrong → accept the first 3, redo from the mistake
→ on average each forward pass "earns" 2–3 tokens for free4. Quantization: The Bandwidth Fix
Every decode step streams the full set of weights through memory, so weight precision directly determines speed: FP16 → INT4 cuts weight traffic by 75%, with near-linear gains in generation speed. For quantization methods and calibration discipline, see Quantization; the key methods are AWQ/GPTQ/GGUF.
5. Prefix Caching (Prompt Caching)
LLM prompts often carry long, unchanging blocks of system instructions and context. Cache the KV results of the shared prefix, and every request that hits the cache skips prefill — TTFT can drop by 50% or more. Multi-turn conversations benefit the most. vLLM's prefix caching and OpenAI's prompt caching are both this idea.
6. Prefill/Decode Disaggregation (PD)
Prefill (compute-bound — it craves FLOPs) and decode (memory-bound — it craves bandwidth) stress hardware differently. Dedicate one pool of GPUs to prefill and another to decode, with scheduling in between, and both resource types can run at full utilization. This suits very large scale and high concurrency; it's an advanced technique.
7. Memory Offloading
Weights and activations move between GPU and CPU memory on demand, trading roughly an order of magnitude in speed for the ability to run a large model on a small GPU — only suitable for "it just needs to run" scenarios (e.g., local tools).
4. Parallelism: Tensor / Pipeline / Data
| Scheme | What gets split | Communication frequency | Best for |
|---|---|---|---|
| Tensor parallel (TP) | Each layer's matrices across GPUs | Multiple times per layer (very high) | Model too big for one GPU, 2–8 GPUs |
| Pipeline parallel (PP) | Layers grouped onto GPUs | Once per stage (low) | Very deep models, cross-node |
| Data parallel (DP) | Full model × N GPUs | Only at synchronization | Throughput scaling, combined with TP/PP |
Engineering note: TP is communication-hungry and needs NVLink/InfiniBand-class interconnects; across nodes, prefer PP over TP. For a deeper treatment, see Papers: Parallel and Distributed Inference.
text
Typical 70B deployment on 8×H100:
TP=8 (tensor split within one node, fully connected over NVLink)
or TP=4 × PP=2 (2 nodes, low inter-node communication)
Let load tests of your actual communication bandwidth make the final call5. Serving Frameworks at a Glance
| Framework | Positioning | Highlights |
|---|---|---|
| vLLM | High-throughput LLM serving | PagedAttention + continuous batching + prefix caching |
| Hugging Face TGI | Deep HF ecosystem integration | Easy to deploy, feature-complete |
| SGLang | Extreme scheduling (RadixAttention) | Strongest performance on complex prompt workloads |
| llama.cpp | Local/CPU friendly | Cross-platform, GGUF ecosystem |
| TensorRT-LLM | The limit of NVIDIA hardware | Graph compilation + a full set of integrated optimizations |
How to choose: balance ecosystem against performance. vLLM is currently the default choice with the widest production adoption; for hands-on experience, see LLM Serving with vLLM.
6. Performance Metrics and Cost Optimization
- TTFT: set by prefill speed; affected by input length and prefill compute;
- TPOT: set by decode speed; affected by weight precision and memory bandwidth;
- Throughput (tokens/s): the direct beneficiary of continuous batching + quantization.
Cost optimization priority order:
text
① Continuous batching (2–5× throughput, nearly free)
② Quantize to INT8/INT4 (halve weight traffic / cut it by 75%)
③ KV cache quantization + GQA (more concurrency / longer context)
④ Prefix caching (big TTFT drop in conversational workloads)
⑤ Speculative decoding (when you have QPS headroom)
⑥ Elastic scaling + PD disaggregation (at larger scale)7. A Typical Deployment Architecture
text
Clients
│ streaming SSE/WebSocket
▼
API gateway (auth / rate limiting / routing)
▼
LLM inference cluster (K8s)
├─ N × vLLM instances (each can load multiple models, routed by model)
│ └─ tensor parallelism (multi-GPU)
▼
Monitoring: TTFT/TPOT/throughput/GPU/memory (see /concepts/monitoring)Trade-offs
| Decision | Options | How to choose |
|---|---|---|
| Throughput vs. latency | Large batch vs. small batch | Protect TTFT for interactive traffic; chase throughput for offline generation |
| Framework | vLLM vs. TensorRT-LLM vs. llama.cpp | vLLM for production at high concurrency; TensorRT-LLM to max out NVIDIA hardware; llama.cpp for local |
| Quantization precision | FP16 vs. INT8 vs. INT4 | Drop precision only when bandwidth-bound, and always validate quality |
| Context length | Longer means more KV cache | Cap it at what the business actually needs; don't chase 128K blindly |
| Single GPU vs. multi-GPU | Simplicity vs. capacity/throughput | Quantize first → then TP → then PP |
In one sentence: LLM inference engineering boils down to balancing three ledgers — the memory ledger (the KV cache is the star), the scheduling ledger (continuous batching keeps every moment busy), and the communication ledger (don't let interconnects strangle your parallelism). Balance all three, and throughput and cost fall into place.
Further Reading
- Papers: PagedAttention — the systems paper on paged KV cache management
- LLM Serving with vLLM — hands-on vLLM, from deployment to tuning
- Quantization — deep dive into GPTQ/AWQ/KV cache quantization
- Papers: Parallel and Distributed Inference — paper-level coverage of TP/PP/DP
- Performance Tuning and Capacity Planning — load testing and capacity math for LLM services
- GPUs and Hardware Selection — the memory ledger and multi-GPU choices
References
- Efficient Memory Management for Large Language Model Serving with PagedAttention (the vLLM paper, SOSP 2023)
- Orca: A Distributed Serving System for Transformer-Based Generative Models (where continuous batching originated)
- Fast Inference from Transformers via Speculative Decoding (Leviathan et al.)
- vLLM documentation
- NVIDIA TensorRT-LLM documentation