Appearance
vLLM and PagedAttention
1. Definition: Turning LLM Inference from "Heavy Artillery" into a "Standard Component"
vLLM is the high-performance LLM inference engine open-sourced by UC Berkeley's Sky Computing Lab in June 2023; its paper, Efficient Memory Management for Large Language Model Serving with PagedAttention, was published at SOSP 2023. The problem it solves can be summarized in one number: HuggingFace Transformers' native inference only reaches 10–25% of a GPU's theoretical throughput; vLLM pushes that to 60–90% — a typical 10–24× throughput gain.
To grasp what vLLM means, you first need to understand the bottlenecks of LLM inference:
- The decode phase is memory-bound (see The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis): generating each token requires reading the entire model weights + the historical KV cache once, while performing very few FLOPs.
- The KV cache is huge and managed crudely: a single Llama-2-7B request with a 2048-token sequence needs ~5 GB of KV cache, and traditional frameworks pre-allocate contiguous GPU memory for the "maximum sequence length" — most requests never fill it, wasting 60–80%.
- Batching is hard: requests differ in length and in the number of generation steps; when long and short requests mix, short ones wait on long ones, and GPU utilization stays low.
vLLM solves all three problems at once with one core innovation (PagedAttention) plus two engineering companions (continuous batching + prefix caching). It is now the de facto standard for open-source LLM inference, alongside TensorRT-LLM, SGLang, and TGI.
2. PagedAttention: OS Virtual Memory Thinking, Brought to LLMs
PagedAttention is vLLM's core innovation, with inspiration taken directly from the operating system's paged virtual memory.
The Problem with the Traditional KV Cache
Traditional frameworks pre-allocate a contiguous stretch of GPU memory per request, sized by max_seq_len:
Request A: max 2048 → actual 800 [████████░░░░░░░░░░░░░░] 1248 wasted
Request B: max 2048 → actual 1500 [█████████████████░░░░░] 548 wasted
Request C: max 2048 → actual 600 [██████░░░░░░░░░░░░░░░░] 1448 wasted
Total 60% GPU memory wasted- Internal fragmentation: large pre-allocation, small actual use — severe waste.
- External fragmentation: holes appear as requests come and go, hard to reuse.
- No sharing: KV of identical prompt prefixes (system prompts) cannot be reused; every request recomputes them.
The PagedAttention Solution
Borrowing from OS virtual memory: the KV cache is divided into blocks (block size is usually 16), and logical addresses are mapped to physical blocks through a block table.
Logical KV: [B0][B1][B2][B3] ← per-request "virtual addresses"
│ │ │ │
▼ ▼ ▼ ▼
block table: │ 1| 3| 7| 9 | ← physical block IDs (scattered in the pool)
▼ ▼
Physical pool: [B0][B1][B2][B3][B4][B5][B6][B7][B8][B9]...
▲ ▲ ▲
└── A └── B └── C- On-demand allocation within a request: the 7th block is only requested when the 100th token is generated — no pre-allocation waste.
- Compact physical memory: all blocks from all requests live scattered in one unified pool — no external fragmentation.
- Shareable prefixes: requests with identical prefixes point to the same physical blocks, with copy-on-write (CoW).
Reworking the Attention Computation
Traditional attention computes "Q × K^T" in one shot, with K/V sitting in contiguous memory. PagedAttention instead loads K/V block by block — each block's attention is computed separately and then accumulated. This adds a small amount of kernel scheduling overhead, but buys a leap in memory utilization.
Key structure of the CUDA kernel (pseudocode):
cpp
// Simplified PagedAttention CUDA kernel
__global__ void paged_attention(
float* output, // [seq_len, d_model]
const float* q, // [num_heads, d_head]
const float* key_cache, // [num_blocks, block_size, num_heads, d_head]
const float* value_cache, // [num_blocks, block_size, num_heads, d_head]
const int* block_table, // [max_num_blocks_per_seq]
int context_len, // history length of the current request
int block_size // usually 16
) {
int block_idx = block_table[block_id]; // logical → physical
// compute attention within the physical block
...
}See the PagedAttention deep dive in Classic Papers in Depth.
3. Continuous Batching: Rebuilding the Batch at Every Step
Traditional batching ("static batching") fills a batch, runs it together, and waits for it to finish together. The problem: short requests get stuck behind long ones, and the GPU idles while waiting for the last request to complete.
Continuous batching (also called "iteration-level batching" or "in-flight batching") takes a different approach: the batch is re-formed at every generation step —
Step 0: [A B C D] 4 requests prefill together
Step 1: [A B C D] 4 decode together
Step 2: [A B C D] A finishes and exits
Step 3: [_ B C D E] E enters, filling A's slot
Step 4: [_ B C D E] D finishes and exits
Step 5: [F B C _ E] F enters, C still going
...At every step, new requests can be pulled in from the prefill queue, and finished ones exit immediately. The GPU stays full and throughput is maximized. See Batching and Request Scheduling.
vLLM's scheduler also implements preemption: when memory pressure builds, the KV of some requests is swapped out to CPU memory (recompute or swap) and brought back once memory frees up. This is a key technique in LLM serving.
4. Prefix Caching: Reusing the System Prompt's KV
LLM applications typically follow the structure of a long system prompt + a short user query. For example, a customer-service bot:
system: "You are a customer-support assistant for XX Corp, responsible for answering the following question types..." ← 1500 tokens
user: "When will my order ship?" ← 8 tokensPrefilling that 1500-token system prompt for every request is a huge waste. vLLM's Automatic Prefix Caching:
- The first request stores its KV in an LRU cache during prefill;
- Later requests with the same system prompt reuse it directly — only the user query is prefilled;
- This saves 80–95% of prefill compute.
Effect (typical customer-service scenario):
| Config | Throughput (tokens/s) | First-token latency |
|---|---|---|
| Prefix caching off | 1200 | 220 ms |
| Prefix caching on | 8500 | 25 ms |
Prefix caching has been enabled by default since vLLM 0.5+.
5. Quantization Support: AWQ / GPTQ / FP8 / INT8 / INT4
vLLM supports mainstream weight-quantization schemes; see Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision:
| Quantization method | Bit width | vLLM support | Notes |
|---|---|---|---|
| GPTQ | INT4 | ✅ Native | Accelerated by EXL2 / Marlin kernels |
| AWQ | INT4 | ✅ Native | Typically 0.3–0.5% better accuracy than GPTQ |
| Marlin (GPTQ-optimized) | INT4 | ✅ | Extreme GPU kernel, 1.5–2× faster |
| bitsandbytes | INT8 / NF4 | ✅ | Good compatibility, mediocre performance |
| FP8 (E4M3) | FP8 | ✅ (H100+) | Hopper's killer feature; see TensorRT-LLM |
| INT8 (W8A16) | INT8 | ✅ | Fallback for older cards |
| INT4 (W4A16) | INT4 | ✅ | Llama-2-7B runs on a single 8 GB GPU |
bash
# Launch an AWQ-quantized model
python -m vllm.entrypoints.openai.api_server \
--model TheBloke/Llama-2-7B-Chat-AWQ \
--quantization awq_marlin \
--dtype float16 \
--max-model-len 4096 \
--gpu-memory-utilization 0.96. Version Evolution: From 0.2 to 0.7+
vLLM evolves at high speed, with every release adding features:
| Version | Date | Key features |
|---|---|---|
| v0.2 | 2023.08 | First release of PagedAttention + continuous batching |
| v0.3 | 2023.11 | AWQ / GPTQ quantization support |
| v0.4 | 2024.02 | Prefix caching, tensor parallel improvements |
| v0.5 | 2024.06 | Speculative decoding (EAGLE/lookahead), chunked prefill, prefix caching on by default |
| v0.6 | 2024.10 | CUDA graphs, FP8 quantization, improved scheduler |
| v0.7+ | 2025.01+ | MTP / EAGLE-3 integration, enhanced multimodality, reasoning-model support |
Key engineering optimizations
Chunked prefill: splits the prefill of long prompts into chunks and schedules them together with decode, so a long prompt cannot starve decode. CUDA graphs: records the entire kernel sequence of a generation step into a graph, cutting launch overhead from a dozen-plus launches down to one. These two brought vLLM's latency down another 30–50% after 0.6+.
7. Performance Data: Baseline Reference
Here is a set of baselines on H100 80GB (vLLM 0.6.x; methodology in Inference Benchmarking in Practice):
| Model | Precision | Concurrency | Throughput (tokens/s) | First-token latency |
|---|---|---|---|---|
| Llama-3-8B | BF16 | 32 | ~5000 | 30 ms |
| Llama-3-8B | FP8 | 32 | ~8500 | 30 ms |
| Llama-3-8B | AWQ-INT4 | 32 | ~7500 | 32 ms |
| Llama-3-70B | BF16 (2×H100 TP) | 32 | ~1500 | 80 ms |
| Llama-3-70B | FP8 (1×H100) | 32 | ~1700 | 75 ms |
| Qwen2-72B | AWQ-INT4 | 32 | ~2000 | 75 ms |
| DeepSeek-R1-Distill-7B | BF16 | 32 | ~5500 | 30 ms |
Against HuggingFace Transformers (FP16, batch=8):
Llama-3-8B HF Transformers: ~280 tokens/s
Llama-3-8B vLLM (BF16): ~5000 tokens/s (18× speedup)
Llama-3-8B vLLM (FP8): ~8500 tokens/s (30× speedup)8. Comparison with SGLang / TGI / TensorRT-LLM
| Engine | Main strengths | Relationship with vLLM |
|---|---|---|
| vLLM | PagedAttention, the broadest community ecosystem, HuggingFace models out of the box | The baseline |
| SGLang | Also from UC Berkeley (2024), RadixAttention prefix tree + structured output + programmatic orchestration | Sibling rivalry; SGLang is stronger for complex applications (structured output, agents, tree search), vLLM is broader for basic serving |
| TGI (HuggingFace) | Official HF offering, deeply integrated with the HF ecosystem | Somewhat weaker performance (~70–85% of vLLM), but lower integration cost within HF |
| TensorRT-LLM | NVIDIA official, extreme optimization on H100 | TRT-LLM is slightly faster on H100 (1.1–1.3×) with a heavier build process; vLLM wins on ease of use |
| DeepSpeed-MII | From Microsoft, ZeRO-Inference supports beyond-memory models | Research-oriented, smaller ecosystem |
| MLC-LLM | General compiler path, cross-platform | Strength in WebGPU / mobile (see Mobile Deployment) |
Choosing an engine
- Getting started, PoCs, small teams → vLLM (ecosystem, ease of use, best docs)
- Extreme production on H100 clusters → TensorRT-LLM or vLLM 0.7+ with CUDA graphs; let benchmarks decide
- Structured output / agents / tree search → SGLang
- Deep HF ecosystem integration → TGI
- Cross-platform / WebGPU / mobile → MLC-LLM
9. Deployment Example: OpenAI-Compatible API in One Line
bash
# Simplest case: start a vLLM OpenAI-compatible server on a single GPU
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-8B-Instruct \
--port 8000 \
--max-model-len 8192 \
--gpu-memory-utilization 0.9 \
--enforce-eager # disable CUDA graphs (for debugging)
# Multi-GPU tensor parallelism (TP=2)
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-70B-Instruct \
--tensor-parallel-size 2 \
--max-model-len 4096
# Call it (OpenAI-compatible protocol)
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Meta-Llama-3-8B-Instruct",
"messages": [{"role": "user", "content": "What is PagedAttention?"}]
}'python
# Python API
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct",
tensor_parallel_size=1,
max_model_len=8192,
quantization=None) # None / "awq" / "gptq" / "fp8"
prompts = ["Explain PagedAttention in one sentence", "What is continuous batching?"]
sampling = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=128)
outputs = llm.generate(prompts, sampling)
for o in outputs:
print(o.outputs[0].text)10. Limitations and Boundaries
- Performance stops growing at very large batches: the decode phase is memory-bound; past a certain concurrency, GPU memory bandwidth saturates and throughput plateaus. At that point, bandwidth-saving techniques like speculative decoding (see Speculative Decoding and Medusa/EAGLE) work better.
- Lagging support for custom models: new architectures (e.g., Mamba, new MoE variants) land on vLLM weeks to months behind HF Transformers.
- Structured output weaker than SGLang: vLLM's guided decoding (outlines / xgrammar) trails SGLang's RadixAttention + complex programmatic orchestration in both performance and flexibility.
- Poor NPU / CPU performance: vLLM is GPU-first; its CPU / NPU backends are weak. For mobile, go with llama.cpp or ONNX Runtime.
- Multimodal support still evolving: vLLM 0.5+ supports image inputs (LLaVA, Qwen-VL), but the ecosystem and optimization level lag plain-text models.
- TP capped at 8: single-machine TP within 8 GPUs; across machines, combine with PP / DP — see Distributed Inference (TP/PP).
11. Where to Go Next
- Concept pages: Batching and Request Scheduling, The GPU Memory Hierarchy and the Bandwidth Wall, The Roofline Model and Compute Analysis, Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision, Model Serving and Orchestration, GPU Architecture and Optimization
- Case-study pages: TensorRT-LLM, Speculative Decoding and Medusa/EAGLE, Triton Inference Server, Distributed Inference (TP/PP)
- Papers: Classic Papers in Depth (PagedAttention deep dive), Frontier Advances, Reading Paths
- Practice pages: Inference Engine Comparison, Tuning and Performance Optimization, Inference Benchmarking in Practice, Common Pitfalls and Anti-Patterns, Deploy an Inference Service from Scratch
- Resource pages: Hardware Primer, Benchmark Data & Tool Profiles, Curated Resources
References
- Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP 2023) — the PagedAttention paper
- vLLM GitHub — main repository
- vLLM documentation — API and configuration reference
- Frantar et al. GPTQ: Accurate Post-Training Quantization (ICLR 2023) — GPTQ quantization
- Lin et al. AWQ: Activation-aware Weight Quantization (MLSys 2024) — AWQ quantization
- Cai et al. Medusa: Simple LLM Inference Acceleration (ICML 2024) — speculative decoding
- Li et al. EAGLE-2: Faster Inference of LLMs with Dynamic Draft Trees (2024) — EAGLE speculative decoding
- Zheng et al. SGLang: Efficient Execution of Structured Language Model Programs (NeurIPS 2024) — SGLang