Skip to content

vLLM and PagedAttention

At a glance The LLM inference engine open-sourced by UC Berkeley in 2023. PagedAttention pages the KV cache like virtual memory, and with continuous batching it lifts Llama-65B throughput by an order of magnitude. This article dissects PagedAttention, prefix caching, quantization support, and the comparison with SGLang/TGI.

vLLM and PagedAttention ​

1. Definition: Turning LLM Inference from "Heavy Artillery" into a "Standard Component" ​

vLLM is the high-performance LLM inference engine open-sourced by UC Berkeley's Sky Computing Lab in June 2023; its paper, Efficient Memory Management for Large Language Model Serving with PagedAttention, was published at SOSP 2023. The problem it solves can be summarized in one number: HuggingFace Transformers' native inference only reaches 10–25% of a GPU's theoretical throughput; vLLM pushes that to 60–90% — a typical 10–24× throughput gain.

To grasp what vLLM means, you first need to understand the bottlenecks of LLM inference:

  • The decode phase is memory-bound (see The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis): generating each token requires reading the entire model weights + the historical KV cache once, while performing very few FLOPs.
  • The KV cache is huge and managed crudely: a single Llama-2-7B request with a 2048-token sequence needs ~5 GB of KV cache, and traditional frameworks pre-allocate contiguous GPU memory for the "maximum sequence length" — most requests never fill it, wasting 60–80%.
  • Batching is hard: requests differ in length and in the number of generation steps; when long and short requests mix, short ones wait on long ones, and GPU utilization stays low.

vLLM solves all three problems at once with one core innovation (PagedAttention) plus two engineering companions (continuous batching + prefix caching). It is now the de facto standard for open-source LLM inference, alongside TensorRT-LLM, SGLang, and TGI.

2. PagedAttention: OS Virtual Memory Thinking, Brought to LLMs ​

PagedAttention is vLLM's core innovation, with inspiration taken directly from the operating system's paged virtual memory.

The Problem with the Traditional KV Cache ​

Traditional frameworks pre-allocate a contiguous stretch of GPU memory per request, sized by max_seq_len:

Request A: max 2048 → actual 800   [████████░░░░░░░░░░░░░░] 1248 wasted
Request B: max 2048 → actual 1500  [█████████████████░░░░░] 548 wasted
Request C: max 2048 → actual 600   [██████░░░░░░░░░░░░░░░░] 1448 wasted
                                                   Total 60% GPU memory wasted
  • Internal fragmentation: large pre-allocation, small actual use — severe waste.
  • External fragmentation: holes appear as requests come and go, hard to reuse.
  • No sharing: KV of identical prompt prefixes (system prompts) cannot be reused; every request recomputes them.

The PagedAttention Solution ​

Borrowing from OS virtual memory: the KV cache is divided into blocks (block size is usually 16), and logical addresses are mapped to physical blocks through a block table.

Logical KV:  [B0][B1][B2][B3]      ← per-request "virtual addresses"
                 │  │  │  │
                 ▼  ▼  ▼  ▼
block table:  │ 1| 3| 7| 9 |        ← physical block IDs (scattered in the pool)
                 ▼        ▼
Physical pool: [B0][B1][B2][B3][B4][B5][B6][B7][B8][B9]...
                     ▲            ▲           ▲
                     └── A        └── B       └── C
  • On-demand allocation within a request: the 7th block is only requested when the 100th token is generated — no pre-allocation waste.
  • Compact physical memory: all blocks from all requests live scattered in one unified pool — no external fragmentation.
  • Shareable prefixes: requests with identical prefixes point to the same physical blocks, with copy-on-write (CoW).

Reworking the Attention Computation ​

Traditional attention computes "Q × K^T" in one shot, with K/V sitting in contiguous memory. PagedAttention instead loads K/V block by block — each block's attention is computed separately and then accumulated. This adds a small amount of kernel scheduling overhead, but buys a leap in memory utilization.

Key structure of the CUDA kernel (pseudocode):

cpp
// Simplified PagedAttention CUDA kernel
__global__ void paged_attention(
    float* output,            // [seq_len, d_model]
    const float* q,           // [num_heads, d_head]
    const float* key_cache,   // [num_blocks, block_size, num_heads, d_head]
    const float* value_cache, // [num_blocks, block_size, num_heads, d_head]
    const int* block_table,   // [max_num_blocks_per_seq]
    int context_len,          // history length of the current request
    int block_size            // usually 16
) {
    int block_idx = block_table[block_id];   // logical → physical
    // compute attention within the physical block
    ...
}

See the PagedAttention deep dive in Classic Papers in Depth.

3. Continuous Batching: Rebuilding the Batch at Every Step ​

Traditional batching ("static batching") fills a batch, runs it together, and waits for it to finish together. The problem: short requests get stuck behind long ones, and the GPU idles while waiting for the last request to complete.

Continuous batching (also called "iteration-level batching" or "in-flight batching") takes a different approach: the batch is re-formed at every generation step —

Step 0:  [A B C D]      4 requests prefill together
Step 1:  [A B C D]      4 decode together
Step 2:  [A B C D]      A finishes and exits
Step 3:  [_ B C D E]    E enters, filling A's slot
Step 4:  [_ B C D E]    D finishes and exits
Step 5:  [F B C _ E]    F enters, C still going
...

At every step, new requests can be pulled in from the prefill queue, and finished ones exit immediately. The GPU stays full and throughput is maximized. See Batching and Request Scheduling.

vLLM's scheduler also implements preemption: when memory pressure builds, the KV of some requests is swapped out to CPU memory (recompute or swap) and brought back once memory frees up. This is a key technique in LLM serving.

4. Prefix Caching: Reusing the System Prompt's KV ​

LLM applications typically follow the structure of a long system prompt + a short user query. For example, a customer-service bot:

system: "You are a customer-support assistant for XX Corp, responsible for answering the following question types..."  ← 1500 tokens
user:   "When will my order ship?"                                                                                     ← 8 tokens

Prefilling that 1500-token system prompt for every request is a huge waste. vLLM's Automatic Prefix Caching:

  1. The first request stores its KV in an LRU cache during prefill;
  2. Later requests with the same system prompt reuse it directly — only the user query is prefilled;
  3. This saves 80–95% of prefill compute.

Effect (typical customer-service scenario):

ConfigThroughput (tokens/s)First-token latency
Prefix caching off1200220 ms
Prefix caching on850025 ms

Prefix caching has been enabled by default since vLLM 0.5+.

5. Quantization Support: AWQ / GPTQ / FP8 / INT8 / INT4 ​

vLLM supports mainstream weight-quantization schemes; see Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision:

Quantization methodBit widthvLLM supportNotes
GPTQINT4✅ NativeAccelerated by EXL2 / Marlin kernels
AWQINT4✅ NativeTypically 0.3–0.5% better accuracy than GPTQ
Marlin (GPTQ-optimized)INT4✅Extreme GPU kernel, 1.5–2× faster
bitsandbytesINT8 / NF4✅Good compatibility, mediocre performance
FP8 (E4M3)FP8✅ (H100+)Hopper's killer feature; see TensorRT-LLM
INT8 (W8A16)INT8✅Fallback for older cards
INT4 (W4A16)INT4✅Llama-2-7B runs on a single 8 GB GPU
bash
# Launch an AWQ-quantized model
python -m vllm.entrypoints.openai.api_server \
    --model TheBloke/Llama-2-7B-Chat-AWQ \
    --quantization awq_marlin \
    --dtype float16 \
    --max-model-len 4096 \
    --gpu-memory-utilization 0.9

6. Version Evolution: From 0.2 to 0.7+ ​

vLLM evolves at high speed, with every release adding features:

VersionDateKey features
v0.22023.08First release of PagedAttention + continuous batching
v0.32023.11AWQ / GPTQ quantization support
v0.42024.02Prefix caching, tensor parallel improvements
v0.52024.06Speculative decoding (EAGLE/lookahead), chunked prefill, prefix caching on by default
v0.62024.10CUDA graphs, FP8 quantization, improved scheduler
v0.7+2025.01+MTP / EAGLE-3 integration, enhanced multimodality, reasoning-model support

Key engineering optimizations

Chunked prefill: splits the prefill of long prompts into chunks and schedules them together with decode, so a long prompt cannot starve decode. CUDA graphs: records the entire kernel sequence of a generation step into a graph, cutting launch overhead from a dozen-plus launches down to one. These two brought vLLM's latency down another 30–50% after 0.6+.

7. Performance Data: Baseline Reference ​

Here is a set of baselines on H100 80GB (vLLM 0.6.x; methodology in Inference Benchmarking in Practice):

ModelPrecisionConcurrencyThroughput (tokens/s)First-token latency
Llama-3-8BBF1632~500030 ms
Llama-3-8BFP832~850030 ms
Llama-3-8BAWQ-INT432~750032 ms
Llama-3-70BBF16 (2×H100 TP)32~150080 ms
Llama-3-70BFP8 (1×H100)32~170075 ms
Qwen2-72BAWQ-INT432~200075 ms
DeepSeek-R1-Distill-7BBF1632~550030 ms

Against HuggingFace Transformers (FP16, batch=8):

Llama-3-8B HF Transformers:  ~280 tokens/s
Llama-3-8B vLLM (BF16):      ~5000 tokens/s    (18× speedup)
Llama-3-8B vLLM (FP8):       ~8500 tokens/s    (30× speedup)

8. Comparison with SGLang / TGI / TensorRT-LLM ​

EngineMain strengthsRelationship with vLLM
vLLMPagedAttention, the broadest community ecosystem, HuggingFace models out of the boxThe baseline
SGLangAlso from UC Berkeley (2024), RadixAttention prefix tree + structured output + programmatic orchestrationSibling rivalry; SGLang is stronger for complex applications (structured output, agents, tree search), vLLM is broader for basic serving
TGI (HuggingFace)Official HF offering, deeply integrated with the HF ecosystemSomewhat weaker performance (~70–85% of vLLM), but lower integration cost within HF
TensorRT-LLMNVIDIA official, extreme optimization on H100TRT-LLM is slightly faster on H100 (1.1–1.3×) with a heavier build process; vLLM wins on ease of use
DeepSpeed-MIIFrom Microsoft, ZeRO-Inference supports beyond-memory modelsResearch-oriented, smaller ecosystem
MLC-LLMGeneral compiler path, cross-platformStrength in WebGPU / mobile (see Mobile Deployment)

Choosing an engine

  • Getting started, PoCs, small teams → vLLM (ecosystem, ease of use, best docs)
  • Extreme production on H100 clusters → TensorRT-LLM or vLLM 0.7+ with CUDA graphs; let benchmarks decide
  • Structured output / agents / tree search → SGLang
  • Deep HF ecosystem integration → TGI
  • Cross-platform / WebGPU / mobile → MLC-LLM

9. Deployment Example: OpenAI-Compatible API in One Line ​

bash
# Simplest case: start a vLLM OpenAI-compatible server on a single GPU
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3-8B-Instruct \
    --port 8000 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.9 \
    --enforce-eager                     # disable CUDA graphs (for debugging)

# Multi-GPU tensor parallelism (TP=2)
python -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Meta-Llama-3-70B-Instruct \
    --tensor-parallel-size 2 \
    --max-model-len 4096

# Call it (OpenAI-compatible protocol)
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B-Instruct",
    "messages": [{"role": "user", "content": "What is PagedAttention?"}]
  }'
python
# Python API
from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct",
          tensor_parallel_size=1,
          max_model_len=8192,
          quantization=None)            # None / "awq" / "gptq" / "fp8"

prompts = ["Explain PagedAttention in one sentence", "What is continuous batching?"]
sampling = SamplingParams(temperature=0.7, top_p=0.9, max_tokens=128)

outputs = llm.generate(prompts, sampling)
for o in outputs:
    print(o.outputs[0].text)

10. Limitations and Boundaries ​

  1. Performance stops growing at very large batches: the decode phase is memory-bound; past a certain concurrency, GPU memory bandwidth saturates and throughput plateaus. At that point, bandwidth-saving techniques like speculative decoding (see Speculative Decoding and Medusa/EAGLE) work better.
  2. Lagging support for custom models: new architectures (e.g., Mamba, new MoE variants) land on vLLM weeks to months behind HF Transformers.
  3. Structured output weaker than SGLang: vLLM's guided decoding (outlines / xgrammar) trails SGLang's RadixAttention + complex programmatic orchestration in both performance and flexibility.
  4. Poor NPU / CPU performance: vLLM is GPU-first; its CPU / NPU backends are weak. For mobile, go with llama.cpp or ONNX Runtime.
  5. Multimodal support still evolving: vLLM 0.5+ supports image inputs (LLaVA, Qwen-VL), but the ecosystem and optimization level lag plain-text models.
  6. TP capped at 8: single-machine TP within 8 GPUs; across machines, combine with PP / DP — see Distributed Inference (TP/PP).

11. Where to Go Next ​

References ​