Skip to content

Tuning and Performance Optimization

At a glance A tuning checklist and performance optimization playbook for LLM inference: detailed coverage of the eight core parameters — max_num_seqs, max_model_len, gpu_memory_utilization, block_size, swap_space, prefix_caching, chunked_prefill, tensor_parallel_size — plus bottleneck diagnosis with nvidia-smi, Nsight Systems, and PyTorch Profiler, and response strategies for compute-bound vs memory-bound workloads.

Tuning and Performance Optimization ​

Tuning is not "trial and error" — it is "measure → look → change": first find which layer the bottleneck lives in, then adjust the parameters for that layer. Blindly tuning 100 parameters just moves the bottleneck somewhere else.

There is a misconception in the inference optimization world: "tuning vLLM means changing max_num_seqs." In fact, vLLM has 8 core parameters, each mapping to a performance dimension, and changing the wrong one just relocates the bottleneck instead of solving it. Changing max_num_seqs won't fix memory-boundedness; changing gpu_memory_utilization won't fix compute-boundedness — this is the canonical application scenario of The Roofline Model and Compute Analysis.

This article provides the complete framework for LLM inference tuning: look at the 8 core parameters first, then walk a performance tuning checklist, and finally use profiling tools to find bottlenecks. After reading it, you should be able to locate which layer any performance problem lives in, pick which parameter to change, and verify the change. This is the tuning companion to Deploy an Inference Service from Scratch and Progressive Tutorial: Three Working Versions.

Scope of this article

This article focuses on vLLM because it is the most mainstream LLM inference engine today. Parameters of SGLang / TensorRT-LLM / TGI are similar in spirit, and the principles transfer. If you want to build the overall performance framework first, read Latency, Throughput, and Concurrency and The GPU Memory Hierarchy and the Bandwidth Wall.

1. General Principle: Measure Before You Tune ​

1. The "Measure → Look → Change" Loop ​

   ┌─────────────────────────────────────────────┐
   │  1 Measure: run [Inference Benchmarking in   │
   │    Practice](/practice/benchmarking)          │
   │    - TTFT, TPOT, p50/p99, memory            │
   └──────────────────┬──────────────────────────┘
                      ▼
   ┌─────────────────────────────────────────────┐
   │  2 Look: profiling                            │
   │    - nvidia-smi (GPU util / memory / temp)   │
   │    - vLLM logs (KV cache usage, throughput)  │
   │    - Nsight Systems (operator-level timing)  │
   └──────────────────┬──────────────────────────┘
                      ▼
   ┌─────────────────────────────────────────────┐
   │  3 Change: locate the bottleneck layer        │
   │    - compute-bound → tune batch/quantization │
   │    - memory-bound → tune quantization/KV     │
   │    - scheduling → tune max_num_seqs/scheduler│
   └──────────────────┬──────────────────────────┘
                      ▼
   ┌─────────────────────────────────────────────┐
   │  4 Verify: measure again, compare versions    │
   │    - Improvement? Move to the next item       │
   │    - No improvement? Roll back, change course │
   └─────────────────────────────────────────────┘

This loop cannot skip steps — jumping straight to "change" without "measure" and "look" is blind tuning, and 9 times out of 10 you change the wrong thing. See "check bottlenecks at all three layers: model, operator, system" in Deployment Design Principles.

2. Tuning Is Not a One-Time Deal ​

No parameter is "tuned once and fixed forever" — the optimal value drifts with business traffic:

  • Daytime peak: raise max_num_seqs (throughput first)
  • Night trough: lower max_num_seqs (latency first)
  • Big-promotion periods: temporarily lower gpu_memory_utilization (squeeze an extra instance onto the GPU)

Production deployment needs the concept of "parameter sets": a peak config, an off-peak config, and a degraded-failover config, switched by scenario. See "parameter sets" in Model Serving and Orchestration.

2. The Eight Core Parameters ​

The 8 parameters below are the entire "main battlefield" of vLLM tuning. For each, we give: meaning, default value, tuning direction, typical impact, and common mistakes.

1. max_num_seqs (Maximum Concurrent Requests) ​

Meaning: The upper bound on requests in flight in vLLM — effectively "the batch size ceiling of continuous batching."

Default: 256

Tuning direction:

ScenarioRecommended valueRationale
Online chat (latency-first)16–32Small batch protects p99
Online batch (throughput-first)64–128Large batch pulls throughput
Offline batch256+Latency irrelevant; maximize throughput
Speculative decoding≤ 16Speculative decoding is a net loss at large batch

Typical impact: Going from 16 → 64 usually triples throughput, but p99 latency can grow 5–10×. This is the core trade-off of Batching and Request Scheduling.

Common mistakes:

  • Too large: 256 ceiling with only 32 concurrent → KV cache memory blows up, some requests get swapped/recomputed, p99 spikes
  • Too small: 64 concurrent against a ceiling of 16 → 48 requests queue, p99 spikes
  • Never adjusted: the same value for peak and trough wastes capacity or degrades

How to find the "inflection point"

Run the "concurrency vs p99 latency" curve from Inference Benchmarking in Practice, find the inflection point where p99 suddenly spikes, and set max_num_seqs just before it. Example:

Concurrency 1  → p99 25 ms
Concurrency 4  → p99 30 ms
Concurrency 16 → p99 50 ms     ← before the inflection point
Concurrency 32 → p99 180 ms    ← inflection point
Concurrency 64 → p99 1200 ms   ← severe degradation

max_num_seqs = 16 is the optimum in this example.

2. max_model_len (Maximum Context Length) ​

Meaning: The maximum number of tokens the model can handle (input + output).

Default: Depends on the model — 4096 for Llama-2-7B

Tuning direction:

ScenarioRecommended valueRationale
Short conversations2048Saves KV cache memory
Long-document summarization8192–16384Business requirement
RAG4096Balance retrieval depth and memory
Ultra-long context (e.g., Claude 200K)32768+Needs large memory + chunked prefill

Typical impact: Going from 4096 → 16384 quadruples KV cache memory and cuts admitted concurrency to 1/4.

Common mistakes:

  • Oversizing "just in case": an 8K workload hard-set to 32K wastes 4× the KV cache memory
  • Ignoring the coupling with max_num_seqs: max_model_len × max_num_seqs is the total KV cache footprint — the two cannot be tuned independently

KV cache memory formula

KV cache memory = 2 × num_layers × hidden_size × num_heads × head_dim
                  × max_model_len × max_num_seqs × dtype_bytes

Llama-2-7B: 32 layers × 128 hidden × 32 heads × 128 head_dim × 4096 × 32 × 2 bytes (bf16) = 8 GB

Set max_num_seqs=64 and it becomes 16 GB. Set max_model_len=16384 and it becomes 32 GB — out of memory.

3. gpu_memory_utilization (GPU Memory Utilization) ​

Meaning: The fraction of total GPU memory vLLM takes (for the KV pool).

Default: 0.9 (90%)

Tuning direction:

ScenarioRecommended valueRationale
Single instance owning the GPU0.85–0.9Maximize the KV pool
Multiple instances sharing the GPU (MPS/cgroup)0.4–0.5Leave room for other instances
Embedder/reranker on the same card0.6–0.7Leave room for other models
Degraded failover (heavy retries)0.5–0.6Headroom for temporary KV allocation

Typical impact: Going from 0.9 → 0.7 shrinks the KV pool by 22% and cuts admitted concurrency by 30%.

Common mistakes:

  • Leaving the default 0.9: trying to run an embedder on the same card OOMs at startup
  • Setting 0.95 to squeeze performance: temporary allocations (kernel intermediates) run out — and it crashes anyway

Don't let vLLM eat the whole GPU

Multi-model coexistence on one GPU is the norm in production. The LLM takes 60–70%; leave 30–40% for the embedder/reranker/monitoring. See "multiple models sharing a GPU without MPS/cgroup" in Common Pitfalls and Anti-Patterns.

4. block_size (KV Cache Block Size) ​

Meaning: The "page size" of PagedAttention. KV cache is managed in blocks of this size, like OS virtual-memory paging.

Default: 16

Tuning direction:

ScenarioRecommended valueRationale
Default16Generally optimal
Extreme latency (small batch)8Reduces internal fragmentation
Extreme throughput (large batch)32Reduces metadata overhead

Typical impact: Going from 16 → 32 halves metadata overhead but increases internal fragmentation. Usually left alone — the vLLM team has heavily optimized for 16.

Common mistakes:

  • Assuming bigger is faster: block_size is not "the bigger the better" — at 64 it actually regresses (internal fragmentation dominates)
  • Changing without verifying: after changing block_size you must re-run benchmarks; never assume

See "PagedAttention implementation" in vLLM and PagedAttention.

5. swap_space (CPU Swap Space) ​

Meaning: When the KV cache fills up, the size of the buffer (GB) for swapping some KV entries out to CPU memory.

Default: 4 (GB)

Tuning direction:

ScenarioRecommended valueRationale
Online serving0Swapping means p99 spikes — better to reject than to swap
Offline batch16–32Occasional swapping is acceptable; memory for throughput
Tight on RAM0Don't eat more when physical memory is short

Typical impact: When a swap-out → swap-in happens, single-request latency jumps from 50 ms to 500 ms+.

Common mistakes:

  • Setting it large as a fallback: assuming swap lets you admit more concurrency — but once swapping starts, p99 collapses
  • Enabling swap on online serving: swap-out/in unpredictability makes p99 uncontrollable

Online serving should disable swap

If an online service triggers swapping regularly, max_num_seqs is set too high — lower the concurrency ceiling instead of enabling swap. Swap is the last safety net, not a routine tool. See "latency-first vs throughput-first: pick one first" in Deployment Design Principles.

6. enable_prefix_caching (Prefix Caching) ​

Meaning: Identical prompt prefixes (e.g., system prompts) are prefilled only once; later requests reuse the KV.

Default: Enabled by default in vLLM 0.5+

Tuning direction:

ScenarioRecommended valueRationale
Multi-turn dialogue / RAG (shared system prompt)OnTTFT drops significantly
Single-turn Q&A (no shared prefix)OffLow hit rate; metadata overhead drags instead
Structured output (shared few-shot)OnFew-shot prefix reuse pays off hugely

Typical impact: In multi-turn dialogue, TTFT drops from 200 ms to 50 ms (prefill skipped outright).

Common mistakes:

  • Not knowing this switch exists: multi-turn scenarios leave it off, TTFT stays high, and you go tune max_num_seqs — treating the wrong symptom
  • Enabling it for single-turn: hit rate < 10%, and the extra metadata overhead slightly raises TPOT

7. enable_chunked_prefill (Chunked Prefill) ​

Meaning: Prefill of long prompts is cut into small chunks that interleave with decode, so a long prompt doesn't block the decode queue.

Default: Enabled by default in vLLM 0.5+

Tuning direction:

ScenarioRecommended valueRationale
Long prompts + online servingOnLong prefills don't block short decodes
Short prompts + offline batchOffChunking adds overhead for no benefit
Ultra-long context (32K+)On + tune max_num_batched_tokensChunking is mandatory or it OOMs

Typical impact: With long prompts, decode-queue p99 drops from 500 ms to 50 ms after enabling chunking.

Common mistakes:

  • Enabling it for offline batch too: batch jobs don't care about decode latency; chunking just adds overhead
  • Not tuning max_num_batched_tokens: the default may not fit your scenario; it needs separate tuning

8. tensor_parallel_size (Tensor Parallelism Degree) ​

Meaning: Across how many GPUs to run tensor parallelism.

Default: 1 (single GPU)

Tuning direction:

ScenarioRecommended valueRationale
7B model1Fits on one card; multiple cards are waste
13B–30B models2Doesn't fit on one card, or KV pool too small
70B model4–8Mandatory
175B+ models8+Mandatory, plus pipeline parallelism

Typical impact: For a 7B model, TP=1 → 2 only gains 1.3× throughput (far from 2×) — communication overhead eats the gains.

Common mistakes:

  • TP for small models too: TP=2 on a 7B model regresses (communication cost > compute saved)
  • TP as a silver bullet: a 70B model still can't be served by TP=8 alone — add PP (pipeline parallelism)

See Distributed Inference (TP/PP).

3. Performance Tuning Checklist ​

Tuning is not picking randomly from the parameter table — walk the steps below in order and you will avoid 90% of blind-tuning traps.

Step 1: nvidia-smi for the Macro View ​

bash
# Continuous monitoring (refresh every second)
nvidia-smi --query-gpu=timestamp,name,utilization.gpu,utilization.memory,memory.used,memory.free,power.draw,temperature.gpu,clocks.sm,clocks.mem --format=csv -l 1

Watch four indicators:

IndicatorHealthy rangeWhat an anomaly means
utilization.gpu70–95%< 50% means idle compute (memory-bound or poor scheduling); sustained > 95% may mean a compute bottleneck
utilization.memory50–90%< 30% means memory bandwidth underused (compute-bound); sustained > 95% may mean memory-bound
temperature.gpu< 80°C> 85°C triggers thermal throttling — 30% performance loss (see "no GPU temperature monitoring" in Common Pitfalls and Anti-Patterns)
power.drawNear TDPFar below TDP means not fully loaded (power wall or scheduling issue)

Step 2: vLLM Logs for the Micro View ​

vLLM prints key information at startup:

INFO 08-22 14:30:00 config.py:65] Initializing KV cache with ... blocks
INFO 08-22 14:30:00 config.py:78] KV cache size: 24576 tokens  # ← watch the KV pool size
INFO 08-22 14:30:00 llm_engine.py:130] # GPU blocks: 1536      # ← number of blocks

Key runtime logs:

# Throughput and latency
INFO 08-22 14:31:00 metrics.py:120] Avg prompt throughput: 250.4 tokens/s
INFO 08-22 14:31:00 metrics.py:121] Avg generation throughput: 1100.2 tokens/s
INFO 08-22 14:31:00 metrics.py:122] Running: 32 reqs (swapped: 0, finished: 28)

Watch for:

  • swapped: N — N > 0 means swapping is happening; lower max_num_seqs or add memory
  • Avg generation throughput — per-instance generation throughput
  • Running: N reqs — requests currently in flight; compare with max_num_seqs to see if it is saturated

Step 3: Nsight Systems for Operator-Level Detail ​

nvidia-smi can't see operator level — for that, Nsight Systems:

bash
# Attach nsys profile when starting vLLM
nsys profile -o vllm_profile \
    --trace=cuda,nvtx,osrt \
    --capture-range=cudaProfilerApi \
    --export=sqlite \
    vllm serve meta-llama/Llama-2-7b-chat-hf --port 8000

Then send a batch of requests, stop vLLM, and analyze the .qdrep file. Look for:

  • Operator time distribution: which kernel dominates? attention / matmul / quantize?
  • CPU-GPU overlap: while the CPU waits for the GPU, is the CPU idle or the GPU waiting on the CPU?
  • Kernel launch overhead: too-frequent launches mean the batch is too small

Step 4: PyTorch Profiler for Framework-Level Detail ​

If you run Hugging Face transformers (not vLLM), PyTorch Profiler is more direct:

python
from torch.profiler import profile, ProfilerActivity

with profile(activities=[ProfilerActivity.CPU, ProfilerActivity.CUDA],
            record_shapes=True) as prof:
    out = model.generate(**inputs, max_new_tokens=64)

print(prof.key_averages().table(sort_by="cuda_time_total", row_limit=20))

Look for:

  • CPU time vs CUDA time: high CPU share → scheduling overhead; high CUDA share → actually computing
  • Operator distribution: how much goes to matmul / attention / elementwise

Step 5: Locate the Bottleneck Layer ​

Combining the four steps above, locate the bottleneck:

ObservationBottleneck layerResponse
GPU util < 50%, memory util < 50%Scheduling / batch too smallRaise max_num_seqs, enable continuous batching
GPU util > 90%, memory util > 90%Compute-boundQuantize (INT8/INT4), add tensor-core optimizations
GPU util < 50%, memory util > 90%Memory-boundQuantize (cut weight traffic), add PagedAttention
GPU temperature > 85°CThermal throttlingImprove cooling, lower max_num_seqs
swapped > 0 sustainedInsufficient memoryLower max_num_seqs, disable swap, add TP
Long-prompt TTFT highSlow prefillEnable chunked prefill, enable prefix caching
High TPOTSlow decodeQuantize, add speculative decoding

This is the engineering application of The Roofline Model and Compute Analysis.

4. Responding to Compute-Bound vs Memory-Bound ​

LLM inference bottlenecks are almost always one of these two. Full response strategies for each follow.

1. Compute-Bound ​

Symptoms:

  • GPU utilization > 90%
  • Memory utilization also high (compute saturated, traffic saturated too)
  • Raising batch no longer increases throughput (compute already saturated)

Root cause: the model itself has too many FLOPs; compute can't keep up.

Responses:

TechniqueGainCostDetails
Quantize (INT8/INT4)2–4× throughputAccuracy lossModel Quantization Fundamentals
Adopt FlashAttention1.3–2×Change attention kernelKernel Fusion and Custom Kernels
Tensor-core optimizations1.5×May need new hardwareGPU Architecture and Optimization
Use a smaller model5–10×Major accuracy dropPruning and Sparsification
Distill a smaller model2–5×Training costKnowledge Distillation

2. Memory-Bound ​

Symptoms:

  • GPU utilization 30–60%
  • Memory utilization > 90%
  • Raising batch doesn't increase throughput (traffic saturated)
  • TPOT far above "compute ÷ tokens"

Root cause: to generate one token, the entire weight must be read from HBM into SRAM — reading is slower than computing.

Responses:

TechniqueGainCostDetails
Weight-only quantization (AWQ/GPTQ)1.5–2× TPOTMinor accuracy lossWeight-Only Quantization and Mixed Precision
PagedAttentionReduces KV wasteNeeds vLLM-class enginevLLM and PagedAttention
Speculative decoding (small batch)2× TPOTExtra draft modelSpeculative Decoding and Medusa/EAGLE
Prefix cachingBig TTFT dropOnly works for multi-turn-
Use a smaller modelMajor bandwidth cutMajor accuracy dropPruning and Sparsification

90% of LLM inference is memory-bound

The decode phase (token generation) is almost always memory-bound — every generated token requires one full pass over the weights, while compute sits far from saturated. That is why quantization pays off far more in LLM scenarios than in classic models — halving weight traffic directly lowers TPOT. This is the special value of Model Quantization Fundamentals in the LLM context.

5. Tuning Case Studies ​

Three real cases showing the tuning workflow.

Case 1: Online Chat p99 Spikes ​

Symptom: vLLM serving Llama-2-7B; at 20 QPS, p99 suddenly jumps from 200 ms to 1200 ms.

Debugging flow:

Step 1: nvidia-smi → GPU util 60%, memory 95%, temperature 75°C (normal)
Step 2: vLLM logs → "swapped: 4 reqs" (swapping happening!)
Step 3: Locate the bottleneck → KV cache insufficient; some requests being swapped out
Step 4: Current config max_num_seqs=64, max_model_len=4096, gpu_mem_util=0.9
Step 5: Change max_num_seqs=32 (before the inflection point), disable swap
Step 6: Verify → swapped=0, p99 back to 180 ms

Key insight: the original config set max_num_seqs too high (64); the KV pool ran short → swap → p99 spiked. Lowering the concurrency ceiling actually lowered p99 — the counterintuitive case from Batching and Request Scheduling.

Case 2: High TTFT on Long Prompts ​

Symptom: vLLM serving Llama-2-7B, prompt 4000 tokens (RAG scenario), TTFT 1.2 s.

Debugging flow:

Step 1: nvidia-smi → normal
Step 2: vLLM logs → "Avg prompt throughput: 6000 tokens/s"
Step 3: Compute → 4000 tokens / 6000 tokens/s = 666 ms prefill
Step 4: The remaining 534 ms is queue wait (short requests blocked by the long prefill)
Step 5: Enable enable_chunked_prefill=true
Step 6: Verify → TTFT down to 380 ms (short requests no longer blocked after chunking)

Key insight: high TTFT on long prompts is not in the prefill itself but in prefill blocking other requests' decode. Chunked prefill interleaves the two, and TTFT drops sharply.

Case 3: Throughput Won't Go Up ​

Symptom: vLLM serving Llama-2-7B, only 800 tok/s at 32 concurrency; theory says 1100+ on an A100.

Debugging flow:

Step 1: nvidia-smi → GPU util 60% (low!), memory 90%
Step 2: vLLM logs → "Running: 32 reqs", no swap
Step 3: Locate → low GPU util + high memory = memory-bound
Step 4: Check the model → bf16 (not quantized)
Step 5: Adopt AWQ INT4 → weight traffic halved
Step 6: Verify → GPU util 85%, throughput 1700 tok/s

Key insight: running an LLM in bf16 at the 7B scale is almost certainly memory-bound — quantization is the first move.

6. Tuning Checklist ​

All tuning disciplines consolidated into one checklist — walk through it before every deployment:

markdown
## Pre-Deployment Tuning Checklist

### 1. Business scenario positioning
- [ ] Latency-first or throughput-first? (determines max_num_seqs)
- [ ] Prompt length distribution? (determines max_model_len)
- [ ] Multi-turn / shared prefix? (determines enable_prefix_caching)

### 2. Hardware budget
- [ ] Memory budget (weights + KV pool + temporaries)
- [ ] Multiple instances sharing the GPU? (determines gpu_memory_utilization)
- [ ] Is cooling adequate? (affects temperature)

### 3. Parameter set
- [ ] max_num_seqs sits before the "concurrency–p99" inflection point
- [ ] max_model_len × max_num_seqs < memory budget
- [ ] gpu_memory_utilization leaves 10% for other processes
- [ ] swap_space = 0 for online serving
- [ ] enable_prefix_caching on for multi-turn scenarios
- [ ] enable_chunked_prefill on for long-prompt scenarios

### 4. Quantization decision
- [ ] Ship bf16, run the business evaluation set
- [ ] Compare against AWQ INT4 + business evaluation set
- [ ] Consider full rollout only if MMLU drops < 1

### 5. Monitoring and alerts
- [ ] nvidia-smi temperature monitoring (alert > 85°C)
- [ ] vLLM logs swap monitoring (alert > 0)
- [ ] Prometheus p99 monitoring (alert beyond 60% of SLA)
- [ ] KV cache hit-rate monitoring (alert < 50%)

### 6. Gray release
- [ ] New config runs on 5% of traffic first
- [ ] Expand to 100% after 24 hours of observation
- [ ] Rollback plan ready (old parameter set)

7. Further Reading ​

References ​