Appearance
LLM Serving with vLLM: OpenAI-Compatible API and Startup Parameter Tuning
In one sentence: vLLM is a high-throughput LLM inference engine built on PagedAttention and continuous batching; a single command deploys an open-source LLM as an OpenAI-compatible HTTP service.
Why it's worth doing: LLM inference is a completely different engineering problem from ordinary model inference — KV cache memory management determines whether you can survive concurrency. The naive approach (say, loading Llama into FastAPI and inferring one request at a time) delivers an order of magnitude lower throughput, because each request monopolizes a chunk of memory and runs serially. vLLM manages the KV cache page by page with PagedAttention (see the PagedAttention paper), and combined with continuous batching, lifts single-GPU throughput from single-digit QPS to hundreds of tokens/s. Using Qwen2.5-7B-Instruct as the example, this article walks through the full deployment, tuning, and load-testing workflow.
1. What vLLM Is and Why It's Fast
Three core techniques:
- PagedAttention: splits the KV cache into fixed-size "pages" allocated on demand, eliminating fragmentation and pre-allocation waste — memory utilization approaches 100% instead of the 60-80% of traditional implementations;
- Continuous batching: traditional batching waits for an entire batch to finish generating before freeing resources; vLLM schedules token by token, so a sequence that finishes immediately makes room for a new request and the GPU never sits idle;
- OpenAI-compatible API: works out of the box with the
openaiPython SDK — zero changes on the business side.
In one line: vLLM is the de facto standard for serving open-source LLMs on one GPU or many. For the underlying KV cache and memory mechanics, see LLM Inference Optimization.
2. Installation and Model Preparation
bash
# Install (Python 3.9+)
pip install vllm # or install the CUDA build matching the official docs
# Check the GPU
nvidia-smi # confirm memory and driver; 7B FP16 weights ≈ 14GB plus KV cache
# The model pulls straight from Hugging Face; the first run auto-downloads to ~/.cache/huggingface
# You can also download it manually beforehand and point to a local path
huggingface-cli download Qwen/Qwen2.5-7B-Instruct --local-dir ./Qwen2.5-7B-InstructThe hardware floor first: Qwen2.5-7B-Instruct FP16 weights are about 14.6GB, so a single 24GB card (A10/4090 class) is the entry-level configuration; for 32B+ models or larger contexts you need multiple GPUs or quantization (see Sections 6 and 7).
3. Launching the OpenAI-Compatible Server
bash
vllm serve Qwen/Qwen2.5-7B-Instruct \
--gpu-memory-utilization 0.9 \
--max-model-len 8192 \
--tensor-parallel-size 1 \
--dtype float16 \
--served-model-name qwen2.5-7b \
--port 8000Once it's up, an OpenAI-compatible API is exposed by default at http://localhost:8000/v1: /v1/chat/completions (chat), /v1/completions (completions), and /v1/models (model list).
The memory estimation formula
KV cache memory available on one GPU ≈ total GPU memory × gpu-memory-utilization − weight memory − activation overhead. On a 24GB card running 7B FP16: 24×0.9 − 14.6 ≈ 7GB of KV cache, enough for roughly 30-50 concurrent sequences at 8192 context (depending on actual token lengths).
4. Key Startup Parameters Explained
| Parameter | Default | What it does | Tuning advice |
|---|---|---|---|
--gpu-memory-utilization | 0.9 | Fraction of GPU memory reserved for vLLM | 0.85-0.95; drop to 0.8 when memory is tight to protect the KV cache |
--max-model-len | model default (e.g. 8192) | Max context length per sequence | Longer wastes memory; set 8K if the business never uses 32K |
--tensor-parallel-size | 1 | Number of GPUs for tensor parallelism | Increase when the model doesn't fit: 2/4/8 |
--dtype | auto | Weight precision | float16 or bfloat16; never float32 (doubles memory) |
--quantization | None | Quantization backend (awq/gptq/…) | See Section 7 on quantized loading |
--enforce-eager | False | Skip CUDA Graph compilation | Enable on slow startups or sporadic compile errors, at a small throughput cost |
--enable-prefix-caching | auto | Reuses KV cache for shared prefixes across requests | Strongly recommend enabling explicitly for chat/Agent workloads |
--served-model-name | model ID | The model name exposed externally | Alias it so business code doesn't change |
--max-num-seqs | 256 | Max concurrent sequences | Lower it when memory is short to avoid excessive queuing |
--max-num-batched-tokens | 8192 | Max tokens processed per step | Reduce it on small GPUs to prevent OOM |
--trust-remote-code | False | Allow loading custom model code | Required by some community models |
The order is "memory budget first, throughput second"
--max-model-len, --max-num-seqs, and --max-num-batched-tokens together determine the memory budget. First check whether the KV cache allotted by --gpu-memory-utilization can support the target concurrency, then tighten parameters one by one — don't max everything out on day one.
Serving Multiple Models in One Instance
vllm serve can host multiple models at once (sharing a single memory pool and KV cache), which suits "light model + heavy model co-location":
bash
vllm serve \
Qwen/Qwen2.5-7B-Instruct \
Qwen/Qwen2.5-0.5B-Instruct \
--served-model-name qwen-7b \
--served-model-name qwen-0.5b \
--gpu-memory-utilization 0.9 \
--max-model-len 8192Clients pick the model with model="qwen-7b" or model="qwen-0.5b". Multiple models on one instance compete for KV cache memory; estimate with the memory formula above before co-locating: 7B + 0.5B weights are about 15.3GB, leaving only ~6GB of shared KV cache for both models on a 24GB card, and concurrency gets diluted. Conclusion first: co-location suits "one service, several traffic lanes", not "both models need high concurrency" — for the latter, split into two service instances or add GPUs.
Working with Orchestration: Containers and GPU Resources
Production environments rarely run vllm serve bare; it usually goes into a container managed by K8s:
bash
docker run --gpus all --shm-size 8g \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model Qwen/Qwen2.5-7B-Instruct \
--gpu-memory-utilization 0.9 \
--max-model-len 8192Watch --shm-size: vLLM relies on shared memory for tokenizer work and data transfer, and the 64MB default triggers /dev/shm out-of-space errors — one of the most common container startup failures. In K8s, request nvidia.com/gpu: 1 for the container and configure liveness/readiness probes against the built-in /health endpoint.
5. Client Calls: the OpenAI SDK and Streaming
python
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1", # vLLM's OpenAI-compatible endpoint
api_key="EMPTY", # vLLM doesn't validate the key by default
)
# Ordinary chat
resp = client.chat.completions.create(
model="qwen2.5-7b",
messages=[{"role": "user", "content": "Explain what a KV cache is in one sentence"}],
max_tokens=256,
temperature=0.7,
)
print(resp.choices[0].message.content)
# Streaming: always use stream for long answers; first-token latency is dramatically lower
stream = client.chat.completions.create(
model="qwen2.5-7b",
messages=[{"role": "user", "content": "Write a poem about model deployment"}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)Why naive FastAPI doesn't fit LLMs
Put an LLM inside FastAPI and infer one request at a time: a single request monopolizes the whole card's memory through a full generation while everything else queues, GPU utilization barely reaches 20%, and you'd have to hand-roll SSE for streaming. vLLM already ships continuous batching, streaming, rate limiting, and metrics — connect the business layer straight through the OpenAI SDK. Bottom line: don't hand-roll a naive service for LLMs; go straight to an inference engine. For the full comparison, see LLM Inference Optimization.
6. Multi-GPU Tensor Parallelism
When one card doesn't fit the model, split it with --tensor-parallel-size:
bash
# Run a 70B quantized model on 2×24GB cards, or larger models on 4
vllm serve Qwen/Qwen2.5-72B-Instruct-AWQ \
--tensor-parallel-size 4 \
--gpu-memory-utilization 0.9 \
--max-model-len 8192Three essentials:
- Inter-GPU communication needs NVLink or direct PCIe; mixing mismatched GPUs drags everyone down to the slowest card;
tensor-parallel-sizemust divide the splittable dimensions of the attention heads/layers — typically 2/4/8;- Tensor parallelism has diminishing returns: 2 GPUs give roughly 1.7-1.9x, 4 GPUs roughly 3-3.5x — not linear. For better price-performance, another route is single-GPU quantization (next section).
7. Loading Quantized Models
Quantization is the key to squeezing a 7B model into less memory and buying higher throughput; see Quantization for the theory. AWQ/GPTQ are the mainstream choices:
bash
# AWQ-quantized: even a 72B model runs on a 24GB card
vllm serve Qwen/Qwen2.5-72B-Instruct-AWQ \
--quantization awq \
--gpu-memory-utilization 0.9 \
--max-model-len 8192
# GPTQ-quantized
vllm serve TheBloke/Llama-2-7B-Chat-GPTQ \
--quantization gptq \
--max-model-len 4096A measured reference for the gains and costs of quantization (7B model, single A10, batched throughput):
| Precision | Weight memory | Relative throughput | Quality impact |
|---|---|---|---|
| FP16 | ~14.6GB | 1.0x (baseline) | None |
| INT8 | ~7.4GB | ~1.2x | Negligible |
| INT4 (AWQ/GPTQ) | ~4.1GB | ~1.5-1.8x | Acceptable for most tasks; slight drop on hard reasoning problems |
Conclusion first: quantize when memory is the bottleneck; stay FP16 when it isn't. INT4 quality loss is usually acceptable, but for tasks involving precise math or multi-step reasoning, run offline evaluations before deciding — don't go by benchmarks alone.
8. Load Testing: Benchmark Tools and Reading the Numbers
vLLM ships its own benchmark scripts; you can also apply the methods from Load Testing and Capacity Planning:
bash
# Simulate 200 concurrent requests, each asked to generate 256 tokens
python benchmarks/benchmark_serving.py \
--model Qwen/Qwen2.5-7B-Instruct \
--tokenizer Qwen/Qwen2.5-7B-Instruct \
--request-rate 200 \
--num-prompts 1000 \
--max-tokens 256 \
--save-result results.jsonThree core metrics to read from the output:
- TTFT (Time To First Token): first-token latency. In streaming scenarios this is what users perceive; target < 500ms-1s;
- TPOT (Time Per Output Token): time per output token, i.e. 1/decode speed — lower means smoother output;
- Throughput (Total token throughput): tokens generated per unit time, distinguishing input and output throughput.
The comparison that best shows vLLM's value is against a naive Hugging Face transformers deployment under the same configuration (7B model, single A10, concurrency 64, 256 output tokens):
| Deployment | Total throughput (token/s) | Per-request generation speed (token/s) | Memory |
|---|---|---|---|
transformers generate, serial | ~300-600 | ~15-30 (single request) | Weights + single-request KV |
| vLLM (continuous batching) | ~3000-5000 | ~17-25 per request | PagedAttention, page-based allocation |
Conclusion first: vLLM's throughput comes from "many requests sharing each decode step"; a single request is no faster than naive deployment — what's faster is the total — so it shines in multi-concurrency scenarios like chat and agents; for single-stream ultra-low latency (one concurrent request, under 15ms per token), consider more specialized optimizations (CUDA Graphs, smaller models).
Reference numbers for 7B/single-GPU (A10 24GB, concurrency 200): TTFT ≈ 300-800ms, TPOT ≈ 40-60ms/token (about 17-25 token/s per request), total throughput ≈ 3000-5000 token/s. The three big throughput levers: raise --gpu-memory-utilization to leave more KV cache, enable --enable-prefix-caching, and quantize.
FAQ and Troubleshooting
| Problem | Symptom | Fix |
|---|---|---|
| OOM at startup | CUDA out of memory / No available memory for the cache blocks | Lower --gpu-memory-utilization, --max-model-len, --max-num-seqs |
| KV cache too small | Logs like CacheConfig: gpu_memory_utilization... reporting insufficient blocks | Give the KV cache more memory; lower the concurrency cap |
| Overlong input | Input ... is longer than max-model-len | Raise --max-model-len (mind the memory) or have clients truncate |
| Concurrency cap | Requests queue up, TTFT spikes | Check --max-num-seqs; raise concurrency or enlarge the KV cache |
| Compile errors / very slow startup | Stuck on CUDA Graph compilation | Add --enforce-eager to skip it; slightly lower throughput but stable |
| Model returns 404 | model not found | --served-model-name doesn't match the name clients call |
| Low prefix hit rate | Identical system prompts recomputed every time | Enable --enable-prefix-caching explicitly, and make the prefix genuinely shared |
For more production pitfalls, see Common Pitfalls and Anti-Patterns.
Further Reading
- PagedAttention: The vLLM Systems Paper — the memory-management theory behind every parameter in this article
- LLM Inference Optimization — a systematic treatment of the KV cache, continuous batching, and parallelism strategies
- Quantization — the theory of AWQ/GPTQ/INT8 and the quality-performance trade-off
- Performance Optimization and Capacity Planning — the metric framework for TTFT/TPOT/throughput and capacity estimation
- Load Testing and Capacity Planning — turning benchmark results into capacity and scaling decisions
- Common Pitfalls and Anti-Patterns — the most frequent LLM-serving pitfalls in production
- Online Serving with FastAPI + Docker — a contrast: naive deployment for ordinary models vs engine-based deployment for LLMs
References
- vLLM documentation: https://docs.vllm.ai/
- vLLM paper (PagedAttention): https://arxiv.org/abs/2309.06180
- Qwen2.5-7B-Instruct model page: https://huggingface.co/Qwen/Qwen2.5-7B-Instruct
- OpenAI Python SDK: https://github.com/openai/openai-python