Appearance
Tuning and Performance Optimization
Tuning is not "trial and error" — it is "measure → look → change": first find which layer the bottleneck lives in, then adjust the parameters for that layer. Blindly tuning 100 parameters just moves the bottleneck somewhere else.
There is a misconception in the inference optimization world: "tuning vLLM means changing max_num_seqs." In fact, vLLM has 8 core parameters, each mapping to a performance dimension, and changing the wrong one just relocates the bottleneck instead of solving it. Changing max_num_seqs won't fix memory-boundedness; changing gpu_memory_utilization won't fix compute-boundedness — this is the canonical application scenario of The Roofline Model and Compute Analysis.
This article provides the complete framework for LLM inference tuning: look at the 8 core parameters first, then walk a performance tuning checklist, and finally use profiling tools to find bottlenecks. After reading it, you should be able to locate which layer any performance problem lives in, pick which parameter to change, and verify the change. This is the tuning companion to Deploy an Inference Service from Scratch and Progressive Tutorial: Three Working Versions.
Scope of this article
This article focuses on vLLM because it is the most mainstream LLM inference engine today. Parameters of SGLang / TensorRT-LLM / TGI are similar in spirit, and the principles transfer. If you want to build the overall performance framework first, read Latency, Throughput, and Concurrency and The GPU Memory Hierarchy and the Bandwidth Wall.
1. General Principle: Measure Before You Tune
1. The "Measure → Look → Change" Loop
┌─────────────────────────────────────────────┐
│ 1 Measure: run [Inference Benchmarking in │
│ Practice](/practice/benchmarking) │
│ - TTFT, TPOT, p50/p99, memory │
└──────────────────┬──────────────────────────┘
▼
┌─────────────────────────────────────────────┐
│ 2 Look: profiling │
│ - nvidia-smi (GPU util / memory / temp) │
│ - vLLM logs (KV cache usage, throughput) │
│ - Nsight Systems (operator-level timing) │
└──────────────────┬──────────────────────────┘
▼
┌─────────────────────────────────────────────┐
│ 3 Change: locate the bottleneck layer │
│ - compute-bound → tune batch/quantization │
│ - memory-bound → tune quantization/KV │
│ - scheduling → tune max_num_seqs/scheduler│
└──────────────────┬──────────────────────────┘
▼
┌─────────────────────────────────────────────┐
│ 4 Verify: measure again, compare versions │
│ - Improvement? Move to the next item │
│ - No improvement? Roll back, change course │
└─────────────────────────────────────────────┘This loop cannot skip steps — jumping straight to "change" without "measure" and "look" is blind tuning, and 9 times out of 10 you change the wrong thing. See "check bottlenecks at all three layers: model, operator, system" in Deployment Design Principles.
2. Tuning Is Not a One-Time Deal
No parameter is "tuned once and fixed forever" — the optimal value drifts with business traffic:
- Daytime peak: raise
max_num_seqs(throughput first) - Night trough: lower
max_num_seqs(latency first) - Big-promotion periods: temporarily lower
gpu_memory_utilization(squeeze an extra instance onto the GPU)
Production deployment needs the concept of "parameter sets": a peak config, an off-peak config, and a degraded-failover config, switched by scenario. See "parameter sets" in Model Serving and Orchestration.
2. The Eight Core Parameters
The 8 parameters below are the entire "main battlefield" of vLLM tuning. For each, we give: meaning, default value, tuning direction, typical impact, and common mistakes.
1. max_num_seqs (Maximum Concurrent Requests)
Meaning: The upper bound on requests in flight in vLLM — effectively "the batch size ceiling of continuous batching."
Default: 256
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| Online chat (latency-first) | 16–32 | Small batch protects p99 |
| Online batch (throughput-first) | 64–128 | Large batch pulls throughput |
| Offline batch | 256+ | Latency irrelevant; maximize throughput |
| Speculative decoding | ≤ 16 | Speculative decoding is a net loss at large batch |
Typical impact: Going from 16 → 64 usually triples throughput, but p99 latency can grow 5–10×. This is the core trade-off of Batching and Request Scheduling.
Common mistakes:
- Too large: 256 ceiling with only 32 concurrent → KV cache memory blows up, some requests get swapped/recomputed, p99 spikes
- Too small: 64 concurrent against a ceiling of 16 → 48 requests queue, p99 spikes
- Never adjusted: the same value for peak and trough wastes capacity or degrades
How to find the "inflection point"
Run the "concurrency vs p99 latency" curve from Inference Benchmarking in Practice, find the inflection point where p99 suddenly spikes, and set max_num_seqs just before it. Example:
Concurrency 1 → p99 25 ms
Concurrency 4 → p99 30 ms
Concurrency 16 → p99 50 ms ← before the inflection point
Concurrency 32 → p99 180 ms ← inflection point
Concurrency 64 → p99 1200 ms ← severe degradationmax_num_seqs = 16 is the optimum in this example.
2. max_model_len (Maximum Context Length)
Meaning: The maximum number of tokens the model can handle (input + output).
Default: Depends on the model — 4096 for Llama-2-7B
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| Short conversations | 2048 | Saves KV cache memory |
| Long-document summarization | 8192–16384 | Business requirement |
| RAG | 4096 | Balance retrieval depth and memory |
| Ultra-long context (e.g., Claude 200K) | 32768+ | Needs large memory + chunked prefill |
Typical impact: Going from 4096 → 16384 quadruples KV cache memory and cuts admitted concurrency to 1/4.
Common mistakes:
- Oversizing "just in case": an 8K workload hard-set to 32K wastes 4× the KV cache memory
- Ignoring the coupling with
max_num_seqs:max_model_len × max_num_seqsis the total KV cache footprint — the two cannot be tuned independently
KV cache memory formula
KV cache memory = 2 × num_layers × hidden_size × num_heads × head_dim
× max_model_len × max_num_seqs × dtype_bytesLlama-2-7B: 32 layers × 128 hidden × 32 heads × 128 head_dim × 4096 × 32 × 2 bytes (bf16) = 8 GB
Set max_num_seqs=64 and it becomes 16 GB. Set max_model_len=16384 and it becomes 32 GB — out of memory.
3. gpu_memory_utilization (GPU Memory Utilization)
Meaning: The fraction of total GPU memory vLLM takes (for the KV pool).
Default: 0.9 (90%)
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| Single instance owning the GPU | 0.85–0.9 | Maximize the KV pool |
| Multiple instances sharing the GPU (MPS/cgroup) | 0.4–0.5 | Leave room for other instances |
| Embedder/reranker on the same card | 0.6–0.7 | Leave room for other models |
| Degraded failover (heavy retries) | 0.5–0.6 | Headroom for temporary KV allocation |
Typical impact: Going from 0.9 → 0.7 shrinks the KV pool by 22% and cuts admitted concurrency by 30%.
Common mistakes:
- Leaving the default 0.9: trying to run an embedder on the same card OOMs at startup
- Setting 0.95 to squeeze performance: temporary allocations (kernel intermediates) run out — and it crashes anyway
Don't let vLLM eat the whole GPU
Multi-model coexistence on one GPU is the norm in production. The LLM takes 60–70%; leave 30–40% for the embedder/reranker/monitoring. See "multiple models sharing a GPU without MPS/cgroup" in Common Pitfalls and Anti-Patterns.
4. block_size (KV Cache Block Size)
Meaning: The "page size" of PagedAttention. KV cache is managed in blocks of this size, like OS virtual-memory paging.
Default: 16
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| Default | 16 | Generally optimal |
| Extreme latency (small batch) | 8 | Reduces internal fragmentation |
| Extreme throughput (large batch) | 32 | Reduces metadata overhead |
Typical impact: Going from 16 → 32 halves metadata overhead but increases internal fragmentation. Usually left alone — the vLLM team has heavily optimized for 16.
Common mistakes:
- Assuming bigger is faster: block_size is not "the bigger the better" — at 64 it actually regresses (internal fragmentation dominates)
- Changing without verifying: after changing block_size you must re-run benchmarks; never assume
See "PagedAttention implementation" in vLLM and PagedAttention.
5. swap_space (CPU Swap Space)
Meaning: When the KV cache fills up, the size of the buffer (GB) for swapping some KV entries out to CPU memory.
Default: 4 (GB)
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| Online serving | 0 | Swapping means p99 spikes — better to reject than to swap |
| Offline batch | 16–32 | Occasional swapping is acceptable; memory for throughput |
| Tight on RAM | 0 | Don't eat more when physical memory is short |
Typical impact: When a swap-out → swap-in happens, single-request latency jumps from 50 ms to 500 ms+.
Common mistakes:
- Setting it large as a fallback: assuming swap lets you admit more concurrency — but once swapping starts, p99 collapses
- Enabling swap on online serving: swap-out/in unpredictability makes p99 uncontrollable
Online serving should disable swap
If an online service triggers swapping regularly, max_num_seqs is set too high — lower the concurrency ceiling instead of enabling swap. Swap is the last safety net, not a routine tool. See "latency-first vs throughput-first: pick one first" in Deployment Design Principles.
6. enable_prefix_caching (Prefix Caching)
Meaning: Identical prompt prefixes (e.g., system prompts) are prefilled only once; later requests reuse the KV.
Default: Enabled by default in vLLM 0.5+
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| Multi-turn dialogue / RAG (shared system prompt) | On | TTFT drops significantly |
| Single-turn Q&A (no shared prefix) | Off | Low hit rate; metadata overhead drags instead |
| Structured output (shared few-shot) | On | Few-shot prefix reuse pays off hugely |
Typical impact: In multi-turn dialogue, TTFT drops from 200 ms to 50 ms (prefill skipped outright).
Common mistakes:
- Not knowing this switch exists: multi-turn scenarios leave it off, TTFT stays high, and you go tune
max_num_seqs— treating the wrong symptom - Enabling it for single-turn: hit rate < 10%, and the extra metadata overhead slightly raises TPOT
7. enable_chunked_prefill (Chunked Prefill)
Meaning: Prefill of long prompts is cut into small chunks that interleave with decode, so a long prompt doesn't block the decode queue.
Default: Enabled by default in vLLM 0.5+
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| Long prompts + online serving | On | Long prefills don't block short decodes |
| Short prompts + offline batch | Off | Chunking adds overhead for no benefit |
| Ultra-long context (32K+) | On + tune max_num_batched_tokens | Chunking is mandatory or it OOMs |
Typical impact: With long prompts, decode-queue p99 drops from 500 ms to 50 ms after enabling chunking.
Common mistakes:
- Enabling it for offline batch too: batch jobs don't care about decode latency; chunking just adds overhead
- Not tuning
max_num_batched_tokens: the default may not fit your scenario; it needs separate tuning
8. tensor_parallel_size (Tensor Parallelism Degree)
Meaning: Across how many GPUs to run tensor parallelism.
Default: 1 (single GPU)
Tuning direction:
| Scenario | Recommended value | Rationale |
|---|---|---|
| 7B model | 1 | Fits on one card; multiple cards are waste |
| 13B–30B models | 2 | Doesn't fit on one card, or KV pool too small |
| 70B model | 4–8 | Mandatory |
| 175B+ models | 8+ | Mandatory, plus pipeline parallelism |
Typical impact: For a 7B model, TP=1 → 2 only gains 1.3× throughput (far from 2×) — communication overhead eats the gains.
Common mistakes:
- TP for small models too: TP=2 on a 7B model regresses (communication cost > compute saved)
- TP as a silver bullet: a 70B model still can't be served by TP=8 alone — add PP (pipeline parallelism)
See Distributed Inference (TP/PP).
3. Performance Tuning Checklist
Tuning is not picking randomly from the parameter table — walk the steps below in order and you will avoid 90% of blind-tuning traps.
Step 1: nvidia-smi for the Macro View
bash
# Continuous monitoring (refresh every second)
nvidia-smi --query-gpu=timestamp,name,utilization.gpu,utilization.memory,memory.used,memory.free,power.draw,temperature.gpu,clocks.sm,clocks.mem --format=csv -l 1Watch four indicators:
| Indicator | Healthy range | What an anomaly means |
|---|---|---|
utilization.gpu | 70–95% | < 50% means idle compute (memory-bound or poor scheduling); sustained > 95% may mean a compute bottleneck |
utilization.memory | 50–90% | < 30% means memory bandwidth underused (compute-bound); sustained > 95% may mean memory-bound |
temperature.gpu | < 80°C | > 85°C triggers thermal throttling — 30% performance loss (see "no GPU temperature monitoring" in Common Pitfalls and Anti-Patterns) |
power.draw | Near TDP | Far below TDP means not fully loaded (power wall or scheduling issue) |
Step 2: vLLM Logs for the Micro View
vLLM prints key information at startup:
INFO 08-22 14:30:00 config.py:65] Initializing KV cache with ... blocks
INFO 08-22 14:30:00 config.py:78] KV cache size: 24576 tokens # ← watch the KV pool size
INFO 08-22 14:30:00 llm_engine.py:130] # GPU blocks: 1536 # ← number of blocksKey runtime logs:
# Throughput and latency
INFO 08-22 14:31:00 metrics.py:120] Avg prompt throughput: 250.4 tokens/s
INFO 08-22 14:31:00 metrics.py:121] Avg generation throughput: 1100.2 tokens/s
INFO 08-22 14:31:00 metrics.py:122] Running: 32 reqs (swapped: 0, finished: 28)Watch for:
swapped: N— N > 0 means swapping is happening; lowermax_num_seqsor add memoryAvg generation throughput— per-instance generation throughputRunning: N reqs— requests currently in flight; compare withmax_num_seqsto see if it is saturated
Step 3: Nsight Systems for Operator-Level Detail
nvidia-smi can't see operator level — for that, Nsight Systems:
bash
# Attach nsys profile when starting vLLM
nsys profile -o vllm_profile \
--trace=cuda,nvtx,osrt \
--capture-range=cudaProfilerApi \
--export=sqlite \
vllm serve meta-llama/Llama-2-7b-chat-hf --port 8000Then send a batch of requests, stop vLLM, and analyze the .qdrep file. Look for:
- Operator time distribution: which kernel dominates? attention / matmul / quantize?
- CPU-GPU overlap: while the CPU waits for the GPU, is the CPU idle or the GPU waiting on the CPU?
- Kernel launch overhead: too-frequent launches mean the batch is too small
Step 4: PyTorch Profiler for Framework-Level Detail
If you run Hugging Face transformers (not vLLM), PyTorch Profiler is more direct:
python
from torch.profiler import profile, ProfilerActivity
with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
record_shapes=True) as prof:
out = model.generate(**inputs, max_new_tokens=64)
print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=20))Look for:
- CPU time vs CUDA time: high CPU share → scheduling overhead; high CUDA share → actually computing
- Operator distribution: how much goes to matmul / attention / elementwise
Step 5: Locate the Bottleneck Layer
Combining the four steps above, locate the bottleneck:
| Observation | Bottleneck layer | Response |
|---|---|---|
| GPU util < 50%, memory util < 50% | Scheduling / batch too small | Raise max_num_seqs, enable continuous batching |
| GPU util > 90%, memory util > 90% | Compute-bound | Quantize (INT8/INT4), add tensor-core optimizations |
| GPU util < 50%, memory util > 90% | Memory-bound | Quantize (cut weight traffic), add PagedAttention |
| GPU temperature > 85°C | Thermal throttling | Improve cooling, lower max_num_seqs |
swapped > 0 sustained | Insufficient memory | Lower max_num_seqs, disable swap, add TP |
| Long-prompt TTFT high | Slow prefill | Enable chunked prefill, enable prefix caching |
| High TPOT | Slow decode | Quantize, add speculative decoding |
This is the engineering application of The Roofline Model and Compute Analysis.
4. Responding to Compute-Bound vs Memory-Bound
LLM inference bottlenecks are almost always one of these two. Full response strategies for each follow.
1. Compute-Bound
Symptoms:
- GPU utilization > 90%
- Memory utilization also high (compute saturated, traffic saturated too)
- Raising batch no longer increases throughput (compute already saturated)
Root cause: the model itself has too many FLOPs; compute can't keep up.
Responses:
| Technique | Gain | Cost | Details |
|---|---|---|---|
| Quantize (INT8/INT4) | 2–4× throughput | Accuracy loss | Model Quantization Fundamentals |
| Adopt FlashAttention | 1.3–2× | Change attention kernel | Kernel Fusion and Custom Kernels |
| Tensor-core optimizations | 1.5× | May need new hardware | GPU Architecture and Optimization |
| Use a smaller model | 5–10× | Major accuracy drop | Pruning and Sparsification |
| Distill a smaller model | 2–5× | Training cost | Knowledge Distillation |
2. Memory-Bound
Symptoms:
- GPU utilization 30–60%
- Memory utilization > 90%
- Raising batch doesn't increase throughput (traffic saturated)
- TPOT far above "compute ÷ tokens"
Root cause: to generate one token, the entire weight must be read from HBM into SRAM — reading is slower than computing.
Responses:
| Technique | Gain | Cost | Details |
|---|---|---|---|
| Weight-only quantization (AWQ/GPTQ) | 1.5–2× TPOT | Minor accuracy loss | Weight-Only Quantization and Mixed Precision |
| PagedAttention | Reduces KV waste | Needs vLLM-class engine | vLLM and PagedAttention |
| Speculative decoding (small batch) | 2× TPOT | Extra draft model | Speculative Decoding and Medusa/EAGLE |
| Prefix caching | Big TTFT drop | Only works for multi-turn | - |
| Use a smaller model | Major bandwidth cut | Major accuracy drop | Pruning and Sparsification |
90% of LLM inference is memory-bound
The decode phase (token generation) is almost always memory-bound — every generated token requires one full pass over the weights, while compute sits far from saturated. That is why quantization pays off far more in LLM scenarios than in classic models — halving weight traffic directly lowers TPOT. This is the special value of Model Quantization Fundamentals in the LLM context.
5. Tuning Case Studies
Three real cases showing the tuning workflow.
Case 1: Online Chat p99 Spikes
Symptom: vLLM serving Llama-2-7B; at 20 QPS, p99 suddenly jumps from 200 ms to 1200 ms.
Debugging flow:
Step 1: nvidia-smi → GPU util 60%, memory 95%, temperature 75°C (normal)
Step 2: vLLM logs → "swapped: 4 reqs" (swapping happening!)
Step 3: Locate the bottleneck → KV cache insufficient; some requests being swapped out
Step 4: Current config max_num_seqs=64, max_model_len=4096, gpu_mem_util=0.9
Step 5: Change max_num_seqs=32 (before the inflection point), disable swap
Step 6: Verify → swapped=0, p99 back to 180 msKey insight: the original config set max_num_seqs too high (64); the KV pool ran short → swap → p99 spiked. Lowering the concurrency ceiling actually lowered p99 — the counterintuitive case from Batching and Request Scheduling.
Case 2: High TTFT on Long Prompts
Symptom: vLLM serving Llama-2-7B, prompt 4000 tokens (RAG scenario), TTFT 1.2 s.
Debugging flow:
Step 1: nvidia-smi → normal
Step 2: vLLM logs → "Avg prompt throughput: 6000 tokens/s"
Step 3: Compute → 4000 tokens / 6000 tokens/s = 666 ms prefill
Step 4: The remaining 534 ms is queue wait (short requests blocked by the long prefill)
Step 5: Enable enable_chunked_prefill=true
Step 6: Verify → TTFT down to 380 ms (short requests no longer blocked after chunking)Key insight: high TTFT on long prompts is not in the prefill itself but in prefill blocking other requests' decode. Chunked prefill interleaves the two, and TTFT drops sharply.
Case 3: Throughput Won't Go Up
Symptom: vLLM serving Llama-2-7B, only 800 tok/s at 32 concurrency; theory says 1100+ on an A100.
Debugging flow:
Step 1: nvidia-smi → GPU util 60% (low!), memory 90%
Step 2: vLLM logs → "Running: 32 reqs", no swap
Step 3: Locate → low GPU util + high memory = memory-bound
Step 4: Check the model → bf16 (not quantized)
Step 5: Adopt AWQ INT4 → weight traffic halved
Step 6: Verify → GPU util 85%, throughput 1700 tok/sKey insight: running an LLM in bf16 at the 7B scale is almost certainly memory-bound — quantization is the first move.
6. Tuning Checklist
All tuning disciplines consolidated into one checklist — walk through it before every deployment:
markdown
## Pre-Deployment Tuning Checklist
### 1. Business scenario positioning
- [ ] Latency-first or throughput-first? (determines max_num_seqs)
- [ ] Prompt length distribution? (determines max_model_len)
- [ ] Multi-turn / shared prefix? (determines enable_prefix_caching)
### 2. Hardware budget
- [ ] Memory budget (weights + KV pool + temporaries)
- [ ] Multiple instances sharing the GPU? (determines gpu_memory_utilization)
- [ ] Is cooling adequate? (affects temperature)
### 3. Parameter set
- [ ] max_num_seqs sits before the "concurrency–p99" inflection point
- [ ] max_model_len × max_num_seqs < memory budget
- [ ] gpu_memory_utilization leaves 10% for other processes
- [ ] swap_space = 0 for online serving
- [ ] enable_prefix_caching on for multi-turn scenarios
- [ ] enable_chunked_prefill on for long-prompt scenarios
### 4. Quantization decision
- [ ] Ship bf16, run the business evaluation set
- [ ] Compare against AWQ INT4 + business evaluation set
- [ ] Consider full rollout only if MMLU drops < 1
### 5. Monitoring and alerts
- [ ] nvidia-smi temperature monitoring (alert > 85°C)
- [ ] vLLM logs swap monitoring (alert > 0)
- [ ] Prometheus p99 monitoring (alert beyond 60% of SLA)
- [ ] KV cache hit-rate monitoring (alert < 50%)
### 6. Gray release
- [ ] New config runs on 5% of traffic first
- [ ] Expand to 100% after 24 hours of observation
- [ ] Rollback plan ready (old parameter set)7. Further Reading
- Latency, Throughput, and Concurrency — the theory of what you tune
- The GPU Memory Hierarchy and the Bandwidth Wall — the root cause of memory-boundedness
- The Roofline Model and Compute Analysis — bottleneck diagnosis method
- Batching and Request Scheduling — the theory behind max_num_seqs
- Model Quantization Fundamentals — the basis of quantization tuning
- Weight-Only Quantization and Mixed Precision — weight-only quantization in detail
- GPU Architecture and Optimization — what to watch on nvidia-smi
- Kernel Fusion and Custom Kernels — what to look for in Nsight Systems
- vLLM and PagedAttention — deep dive into vLLM's implementation
- Inference Benchmarking in Practice — what to measure before tuning
- Deployment Design Principles — the discipline of tuning
- Common Pitfalls and Anti-Patterns — the collection of tuning anti-patterns
- Deploy an Inference Service from Scratch — the end-to-end tuning scenario
References
- vLLM: Configuration — official documentation of all vLLM parameters
- vLLM: Performance Optimization — vLLM performance optimization guide
- NVIDIA Nsight Systems — GPU profiling tool
- PyTorch Profiler — framework-level profiler
- nvidia-smi Documentation — GPU monitoring
- NVIDIA Data Center GPU Specs — specifications of each GPU model
- vLLM: PagedAttention Paper — the PagedAttention paper
- SGLang: RadixAttention — the RadixAttention paper
- Continuous Batching: Orca Paper — the continuous batching paper