Skip to content

LLM Serving with vLLM

At a glance A hands-on guide to deploying open-source LLMs (Qwen/Llama and friends) with vLLM behind an OpenAI-compatible API: startup parameter tuning, loading quantized models, tensor parallelism, load testing, and troubleshooting.

LLM Serving with vLLM: OpenAI-Compatible API and Startup Parameter Tuning ​

In one sentence: vLLM is a high-throughput LLM inference engine built on PagedAttention and continuous batching; a single command deploys an open-source LLM as an OpenAI-compatible HTTP service.

Why it's worth doing: LLM inference is a completely different engineering problem from ordinary model inference — KV cache memory management determines whether you can survive concurrency. The naive approach (say, loading Llama into FastAPI and inferring one request at a time) delivers an order of magnitude lower throughput, because each request monopolizes a chunk of memory and runs serially. vLLM manages the KV cache page by page with PagedAttention (see the PagedAttention paper), and combined with continuous batching, lifts single-GPU throughput from single-digit QPS to hundreds of tokens/s. Using Qwen2.5-7B-Instruct as the example, this article walks through the full deployment, tuning, and load-testing workflow.

1. What vLLM Is and Why It's Fast ​

Three core techniques:

  1. PagedAttention: splits the KV cache into fixed-size "pages" allocated on demand, eliminating fragmentation and pre-allocation waste — memory utilization approaches 100% instead of the 60-80% of traditional implementations;
  2. Continuous batching: traditional batching waits for an entire batch to finish generating before freeing resources; vLLM schedules token by token, so a sequence that finishes immediately makes room for a new request and the GPU never sits idle;
  3. OpenAI-compatible API: works out of the box with the openai Python SDK — zero changes on the business side.

In one line: vLLM is the de facto standard for serving open-source LLMs on one GPU or many. For the underlying KV cache and memory mechanics, see LLM Inference Optimization.

2. Installation and Model Preparation ​

bash
# Install (Python 3.9+)
pip install vllm          # or install the CUDA build matching the official docs

# Check the GPU
nvidia-smi                # confirm memory and driver; 7B FP16 weights ≈ 14GB plus KV cache

# The model pulls straight from Hugging Face; the first run auto-downloads to ~/.cache/huggingface
# You can also download it manually beforehand and point to a local path
huggingface-cli download Qwen/Qwen2.5-7B-Instruct --local-dir ./Qwen2.5-7B-Instruct

The hardware floor first: Qwen2.5-7B-Instruct FP16 weights are about 14.6GB, so a single 24GB card (A10/4090 class) is the entry-level configuration; for 32B+ models or larger contexts you need multiple GPUs or quantization (see Sections 6 and 7).

3. Launching the OpenAI-Compatible Server ​

bash
vllm serve Qwen/Qwen2.5-7B-Instruct \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192 \
  --tensor-parallel-size 1 \
  --dtype float16 \
  --served-model-name qwen2.5-7b \
  --port 8000

Once it's up, an OpenAI-compatible API is exposed by default at http://localhost:8000/v1: /v1/chat/completions (chat), /v1/completions (completions), and /v1/models (model list).

The memory estimation formula

KV cache memory available on one GPU ≈ total GPU memory × gpu-memory-utilization − weight memory − activation overhead. On a 24GB card running 7B FP16: 24×0.9 − 14.6 ≈ 7GB of KV cache, enough for roughly 30-50 concurrent sequences at 8192 context (depending on actual token lengths).

4. Key Startup Parameters Explained ​

ParameterDefaultWhat it doesTuning advice
--gpu-memory-utilization0.9Fraction of GPU memory reserved for vLLM0.85-0.95; drop to 0.8 when memory is tight to protect the KV cache
--max-model-lenmodel default (e.g. 8192)Max context length per sequenceLonger wastes memory; set 8K if the business never uses 32K
--tensor-parallel-size1Number of GPUs for tensor parallelismIncrease when the model doesn't fit: 2/4/8
--dtypeautoWeight precisionfloat16 or bfloat16; never float32 (doubles memory)
--quantizationNoneQuantization backend (awq/gptq/…)See Section 7 on quantized loading
--enforce-eagerFalseSkip CUDA Graph compilationEnable on slow startups or sporadic compile errors, at a small throughput cost
--enable-prefix-cachingautoReuses KV cache for shared prefixes across requestsStrongly recommend enabling explicitly for chat/Agent workloads
--served-model-namemodel IDThe model name exposed externallyAlias it so business code doesn't change
--max-num-seqs256Max concurrent sequencesLower it when memory is short to avoid excessive queuing
--max-num-batched-tokens8192Max tokens processed per stepReduce it on small GPUs to prevent OOM
--trust-remote-codeFalseAllow loading custom model codeRequired by some community models

The order is "memory budget first, throughput second"

--max-model-len, --max-num-seqs, and --max-num-batched-tokens together determine the memory budget. First check whether the KV cache allotted by --gpu-memory-utilization can support the target concurrency, then tighten parameters one by one — don't max everything out on day one.

Serving Multiple Models in One Instance ​

vllm serve can host multiple models at once (sharing a single memory pool and KV cache), which suits "light model + heavy model co-location":

bash
vllm serve \
  Qwen/Qwen2.5-7B-Instruct \
  Qwen/Qwen2.5-0.5B-Instruct \
  --served-model-name qwen-7b \
  --served-model-name qwen-0.5b \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192

Clients pick the model with model="qwen-7b" or model="qwen-0.5b". Multiple models on one instance compete for KV cache memory; estimate with the memory formula above before co-locating: 7B + 0.5B weights are about 15.3GB, leaving only ~6GB of shared KV cache for both models on a 24GB card, and concurrency gets diluted. Conclusion first: co-location suits "one service, several traffic lanes", not "both models need high concurrency" — for the latter, split into two service instances or add GPUs.

Working with Orchestration: Containers and GPU Resources ​

Production environments rarely run vllm serve bare; it usually goes into a container managed by K8s:

bash
docker run --gpus all --shm-size 8g \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -p 8000:8000 \
  vllm/vllm-openai:latest \
  --model Qwen/Qwen2.5-7B-Instruct \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192

Watch --shm-size: vLLM relies on shared memory for tokenizer work and data transfer, and the 64MB default triggers /dev/shm out-of-space errors — one of the most common container startup failures. In K8s, request nvidia.com/gpu: 1 for the container and configure liveness/readiness probes against the built-in /health endpoint.

5. Client Calls: the OpenAI SDK and Streaming ​

python
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",   # vLLM's OpenAI-compatible endpoint
    api_key="EMPTY",                        # vLLM doesn't validate the key by default
)

# Ordinary chat
resp = client.chat.completions.create(
    model="qwen2.5-7b",
    messages=[{"role": "user", "content": "Explain what a KV cache is in one sentence"}],
    max_tokens=256,
    temperature=0.7,
)
print(resp.choices[0].message.content)

# Streaming: always use stream for long answers; first-token latency is dramatically lower
stream = client.chat.completions.create(
    model="qwen2.5-7b",
    messages=[{"role": "user", "content": "Write a poem about model deployment"}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)

Why naive FastAPI doesn't fit LLMs

Put an LLM inside FastAPI and infer one request at a time: a single request monopolizes the whole card's memory through a full generation while everything else queues, GPU utilization barely reaches 20%, and you'd have to hand-roll SSE for streaming. vLLM already ships continuous batching, streaming, rate limiting, and metrics — connect the business layer straight through the OpenAI SDK. Bottom line: don't hand-roll a naive service for LLMs; go straight to an inference engine. For the full comparison, see LLM Inference Optimization.

6. Multi-GPU Tensor Parallelism ​

When one card doesn't fit the model, split it with --tensor-parallel-size:

bash
# Run a 70B quantized model on 2×24GB cards, or larger models on 4
vllm serve Qwen/Qwen2.5-72B-Instruct-AWQ \
  --tensor-parallel-size 4 \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192

Three essentials:

  1. Inter-GPU communication needs NVLink or direct PCIe; mixing mismatched GPUs drags everyone down to the slowest card;
  2. tensor-parallel-size must divide the splittable dimensions of the attention heads/layers — typically 2/4/8;
  3. Tensor parallelism has diminishing returns: 2 GPUs give roughly 1.7-1.9x, 4 GPUs roughly 3-3.5x — not linear. For better price-performance, another route is single-GPU quantization (next section).

7. Loading Quantized Models ​

Quantization is the key to squeezing a 7B model into less memory and buying higher throughput; see Quantization for the theory. AWQ/GPTQ are the mainstream choices:

bash
# AWQ-quantized: even a 72B model runs on a 24GB card
vllm serve Qwen/Qwen2.5-72B-Instruct-AWQ \
  --quantization awq \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192

# GPTQ-quantized
vllm serve TheBloke/Llama-2-7B-Chat-GPTQ \
  --quantization gptq \
  --max-model-len 4096

A measured reference for the gains and costs of quantization (7B model, single A10, batched throughput):

PrecisionWeight memoryRelative throughputQuality impact
FP16~14.6GB1.0x (baseline)None
INT8~7.4GB~1.2xNegligible
INT4 (AWQ/GPTQ)~4.1GB~1.5-1.8xAcceptable for most tasks; slight drop on hard reasoning problems

Conclusion first: quantize when memory is the bottleneck; stay FP16 when it isn't. INT4 quality loss is usually acceptable, but for tasks involving precise math or multi-step reasoning, run offline evaluations before deciding — don't go by benchmarks alone.

8. Load Testing: Benchmark Tools and Reading the Numbers ​

vLLM ships its own benchmark scripts; you can also apply the methods from Load Testing and Capacity Planning:

bash
# Simulate 200 concurrent requests, each asked to generate 256 tokens
python benchmarks/benchmark_serving.py \
  --model Qwen/Qwen2.5-7B-Instruct \
  --tokenizer Qwen/Qwen2.5-7B-Instruct \
  --request-rate 200 \
  --num-prompts 1000 \
  --max-tokens 256 \
  --save-result results.json

Three core metrics to read from the output:

  • TTFT (Time To First Token): first-token latency. In streaming scenarios this is what users perceive; target < 500ms-1s;
  • TPOT (Time Per Output Token): time per output token, i.e. 1/decode speed — lower means smoother output;
  • Throughput (Total token throughput): tokens generated per unit time, distinguishing input and output throughput.

The comparison that best shows vLLM's value is against a naive Hugging Face transformers deployment under the same configuration (7B model, single A10, concurrency 64, 256 output tokens):

DeploymentTotal throughput (token/s)Per-request generation speed (token/s)Memory
transformers generate, serial~300-600~15-30 (single request)Weights + single-request KV
vLLM (continuous batching)~3000-5000~17-25 per requestPagedAttention, page-based allocation

Conclusion first: vLLM's throughput comes from "many requests sharing each decode step"; a single request is no faster than naive deployment — what's faster is the total — so it shines in multi-concurrency scenarios like chat and agents; for single-stream ultra-low latency (one concurrent request, under 15ms per token), consider more specialized optimizations (CUDA Graphs, smaller models).

Reference numbers for 7B/single-GPU (A10 24GB, concurrency 200): TTFT ≈ 300-800ms, TPOT ≈ 40-60ms/token (about 17-25 token/s per request), total throughput ≈ 3000-5000 token/s. The three big throughput levers: raise --gpu-memory-utilization to leave more KV cache, enable --enable-prefix-caching, and quantize.

FAQ and Troubleshooting ​

ProblemSymptomFix
OOM at startupCUDA out of memory / No available memory for the cache blocksLower --gpu-memory-utilization, --max-model-len, --max-num-seqs
KV cache too smallLogs like CacheConfig: gpu_memory_utilization... reporting insufficient blocksGive the KV cache more memory; lower the concurrency cap
Overlong inputInput ... is longer than max-model-lenRaise --max-model-len (mind the memory) or have clients truncate
Concurrency capRequests queue up, TTFT spikesCheck --max-num-seqs; raise concurrency or enlarge the KV cache
Compile errors / very slow startupStuck on CUDA Graph compilationAdd --enforce-eager to skip it; slightly lower throughput but stable
Model returns 404model not found--served-model-name doesn't match the name clients call
Low prefix hit rateIdentical system prompts recomputed every timeEnable --enable-prefix-caching explicitly, and make the prefix genuinely shared

For more production pitfalls, see Common Pitfalls and Anti-Patterns.

Further Reading ​

References ​