Appearance
Inference Benchmarking in Practice
A wrong number is more dangerous than no number at all — it dresses an optimization decision in "science" while hiding the real bottleneck.
There is an old saying in the inference optimization world: "my model runs faster than yours on your hardware" — that is marketing talk, not an engineering conclusion. The same model on the same GPU can produce tokens/s numbers that differ by 3× between two teams, and the reason is the test method: different batch sizes, different prompt lengths, different quantization precision, different KV cache configurations, no cold/warm start separation. A benchmark without methodology is worse than no benchmark — it keeps you investing in the wrong direction.
This article lays out a reusable, comparable, cheat-proof methodology for LLM inference benchmarking. After reading it, you should be able to write a test report that anyone can reproduce for any two-engine comparison, recognize the five most common forms of "benchmark cheating," and turn test results into decision criteria. This is the source methodology for the performance numbers in Deploy an Inference Service from Scratch and Progressive Tutorial: Three Working Versions.
Scope of this article
This article focuses on single-instance LLM inference benchmarking — engine comparisons on a single machine with one or more GPUs. Distributed system-level benchmarking (cluster QPS, cross-instance latency, network overhead) is a separate methodology and is not covered here. See Triton Inference Server and Model Serving and Orchestration.
1. What to Measure: Core Metrics for LLM Inference
The biggest difference between LLM inference and classic ML inference: the metric is not one number but a set. Looking only at tokens/s is like looking only at an average score — it hides long-tail latency, hides time to first token, and hides concurrency degradation. The five metrics below are all mandatory.
1. The Five Core Metrics
| Metric | Definition | Business meaning | What it measures |
|---|---|---|---|
| TTFT (Time To First Token) | Time from request to first received token | "Responsiveness" in chat scenarios; the boundary of user patience | Prefill phase + queue wait |
| TPOT (Time Per Output Token) | Average per-token time during generation | "Typing speed" of streaming generation | Decode phase + KV cache access |
| End-to-end latency (E2E) | Time from request to complete output | "Wait-for-result" scenarios: batch jobs, document summarization | TTFT + N × TPOT |
| Throughput (tokens/s) | Total tokens generated per unit time | Cluster throughput capacity, cost accounting | (Concurrency × output_len) / total time |
| Peak GPU memory | Maximum memory occupied during inference | Deployment feasibility, concurrency ceiling | Weights + KV pool + temporaries |
2. Latency Distribution: p50 / p95 / p99
Average latency lies — 28 of 30 concurrent requests take 1 s and 2 take 10 s, so the average of 1.6 s looks fine, but those two 10-second requests have already lost the user in a chat scenario. Production must look at percentiles:
Latency distribution (32 concurrency, sorted by TPOT):
p50 = 22 ms ← median, "the normal-case experience"
p90 = 35 ms ← 90% of requests are faster than this
p95 = 50 ms ← the "long tail" SREs care about
p99 = 180 ms ← the truly "unlucky users"
p999 = 1200 ms ← extreme cases when the system is degrading- p50: representative of user experience, but hides the long tail
- p95: the SRE's primary view — it represents "the floor most users get"
- p99: watch this in cost-sensitive scenarios — pushing p99 from 200 ms to 180 ms may cost double the machines
- p999: usually not reported, because the extreme 0.1% cases tend to come from GC, network jitter, and other non-model factors
Discipline on averages vs. percentiles
- External reporting: report p50 + p99 (a pair reflecting center and tail)
- Internal SRE: report p50/p90/p95/p99/p999 (a set reflecting the shape of the distribution)
- Never report only the average — one slow request gets averaged away by 99 fast ones, and the problem disappears from view
See "latency distribution and queuing theory" in Latency, Throughput, and Concurrency.
3. The Three Layers of Peak Memory
Memory is not one number; it is three layers:
| Layer | Contents | Optimized by |
|---|---|---|
| Weights | The model parameters themselves | Quantization, pruning, distillation (see Model Quantization Fundamentals) |
| KV cache | K/V tensors of past tokens | PagedAttention, block_size, max_num_seqs |
| Temporaries | Attention intermediate matrices, activations | Kernel fusion, FlashAttention |
The "memory usage" you see in nvidia-smi is the sum of all three — without looking at each layer, you are tuning blind: "Why does a 14 GB model occupy 36 GB?" Because the KV pool takes 22 GB. See The GPU Memory Hierarchy and the Bandwidth Wall.
2. What to Measure With: A Tour of Tools
Tools for LLM inference benchmarking fall into four categories; each measures different metrics and serves different purposes.
1. Tool Matrix
| Tool | Primary use | What it measures | Maintainer |
|---|---|---|---|
vLLM benchmarks/ | vLLM's own performance benchmarking and comparisons | Throughput, latency, memory | vLLM project |
| MLPerf-LLM | Standardized benchmark across hardware and engines | Throughput, latency, cost | MLCommons |
SGLang bench/ | SGLang vs. vLLM comparisons | Throughput, latency, complex scheduling | SGLang project |
| lm-eval-harness | Model quality evaluation (not performance) | MMLU/HellaSwag and other accuracy metrics | EleutherAI |
| trtllm-bench | TensorRT-LLM performance benchmarking | Throughput, latency | NVIDIA |
| openllm-bench | Cross-engine, cross-quantization comparison | All dimensions | Community |
2. vLLM Offline Inference Benchmark (Most Common)
vLLM's bundled benchmarks/benchmark_throughput.py and benchmark_latency.py are the most common "entry-level" benchmarks:
bash
# Offline throughput benchmark (no HTTP; direct Python API)
python benchmarks/benchmark_throughput.py \
--model meta-llama/Llama-2-7b-chat-hf \
--backend vllm \
--input-len 1024 \
--output-len 256 \
--num-prompts 512 \
--dtype bfloat16 \
--max-model-len 4096
# Single-request latency benchmark
python benchmarks/benchmark_latency.py \
--model meta-llama/Llama-2-7b-chat-hf \
--backend vllm \
--input-len 1024 \
--output-len 256 \
--batch-size 1 \
--num-iters 100The advantage: it talks directly to vLLM's internal API with no HTTP layer overhead, so it measures the engine itself. The downside: it only measures vLLM (other engines need their own scripts). See vLLM and PagedAttention.
3. MLPerf-LLM (Standardized Comparison)
MLPerf-LLM is a standardized benchmark across hardware and engines — all participants test under identical rules, and results are publicly comparable:
| Dimension | What it measures |
|---|---|
| Single-stream | Single-request latency (TTFT + TPOT) |
| Multi-stream | Latency distribution under concurrent requests |
| Offline | Pure throughput (no concurrency constraints) |
| Server | Online QPS and SLO attainment |
MLPerf's discipline is strict: every participant must report the full configuration (hardware model, driver version, engine version, model weight hash), and results must be reproducible. This is the gold standard for industrial-grade comparison — if your vLLM number of 1100 tok/s on an A100 can't be reconciled with the public MLPerf leaderboard, the problem is your method.
4. SGLang Bench (Comparative Testing)
SGLang's bench/ provides comparison scripts for vLLM and SGLang — it covers complex-scheduling scenarios (multi-turn dialogue, long context, structured output) more thoroughly than vLLM's bundled scripts:
bash
# SGLang vs vLLM comparison
python -m sglang.bench_serving \
--backend vllm \
--base-url http://localhost:8000 \
--model meta-llama/Llama-2-7b-chat-hf \
--dataset-name random \
--random-input-len 1024 \
--random-output-len 256 \
--num-prompts 512See vLLM and PagedAttention and "complex scheduling scenarios" in Inference Engine Comparison.
5. lm-eval-harness (Model Quality)
Performance is not everything — after quantization, pruning, or distillation, has the model gotten dumber? That is lm-evaluation-harness's job:
bash
# Run the MMLU 5-shot evaluation
lm_eval --model vllm \
--model_args pretrained=meta-llama/Llama-2-7b-chat-hf \
--tasks mmlu \
--num_fewshot 5 \
--batch_size 8
# Run the AWQ-quantized model
lm_eval --model vllm \
--model_args pretrained=./Llama-2-7b-chat-hf-awq,quantization=awq \
--tasks mmlu,hellaswag,arc_challenge \
--num_fewshot 5Performance numbers must always be reported together with quality numbers — reporting tokens/s without MMLU is like reporting cost without effect. See "quantization is a trade-off" in Model Quantization Fundamentals.
3. Key Variables: What Moves the Numbers Significantly
Each of the variables below can change the numbers by more than 2×. A report must explicitly state the value of every variable, or it cannot be reproduced.
1. Variable Matrix
| Variable | Default | Direction of impact | Typical effect |
|---|---|---|---|
batch_size / max_num_seqs | 1 / 256 | Larger → throughput ↑ but p99 latency ↑ | 5–10× |
input_len (prompt length) | Not fixed | Longer → TTFT ↑, KV cache ↑ | Linear |
output_len (generation length) | Not fixed | Longer → per-request latency ↑, TPOT unchanged | Linear |
| Concurrency | 1 | Higher → throughput ↑ but p99 latency ↑ | 3–30× |
| Quantization precision | bf16 | INT4 → memory ↓ and TPOT ↓ | 1.3–2× |
| KV cache size (gpu_mem_util × block_size) | 0.9 × 16 | Larger → admits more concurrency | 2–5× |
| Model | Not fixed | Larger → every metric degrades | 5–10× |
| GPU model | A100 40GB | Newer is faster (A100→H100: 2×) | 1.5–3× |
| CUDA / driver version | Latest | Mismatch can cost 30% | 1.0–1.5× |
2. Discipline for Fair Comparisons
When comparing two engines, you must hold constant:
Same model (same weight hash)
Same GPU (same model, same driver)
Same batch_size or concurrency
Same input_len and output_len
Same quantization precision
Same max_model_len and KV pool
Single variable: the engine itselfMiss one variable and the comparison is unfair. The classic counterexample: "vLLM has PagedAttention on by default, TGI doesn't" — run both with default settings and vLLM wins by default, but what wins is not the engine, it is the KV pool configuration. A fair comparison must tune both sides to their respective best configurations, not their defaults.
Comparing "default configurations" is a scam
Most "vLLM is 5× faster than TGI" comparison articles are default-configuration comparisons — PagedAttention enabled on one side, disabled on the other. The real comparison is "best-tuned vLLM vs best-tuned TGI" — that is the state users will actually tune the product into. See Inference Engine Comparison.
4. Cheat-Proofing: The 5 Most Common Tricks
The five forms of "cheating" in benchmark circles — most are not intentional, just mis-measurement that went unnoticed:
1. Not Separating Cold Start from Warm Start
Symptom: The first run is slow, later runs are fast, and the report only cites the later ones.
Root cause: CUDA kernel compilation, CUDA Graph capture, and KV cache warmup all happen in the first few inferences. A cold start being 3–10× slower is normal.
Fix:
python
# Standard practice: warmup N times (N ≥ 3), then benchmark
for _ in range(3):
_ = model.generate(**inputs, max_new_tokens=8) # warmup
# Formal test
t0 = time.perf_counter()
out = model.generate(...)
t1 = time.perf_counter()State in the report: "3 warmup runs, then 100 measured runs, p50/p99 reported."
2. Ignoring CPU-GPU Overlap
Symptom: Using time.time() to measure end-to-end, which includes Python interpretation, HTTP serialization, and other CPU time.
Root cause: CUDA is asynchronous — when model.generate returns, the GPU is still running. Timing without torch.cuda.synchronize() measures "the time Python takes to launch the GPU."
Fix:
python
torch.cuda.synchronize()
t0 = time.perf_counter()
out = model.generate(...)
torch.cuda.synchronize() # ← wait for the GPU to actually finish
t1 = time.perf_counter()For HTTP service tests, use the time when the client has received the last chunk, not the time of request.send.
3. Padding Inflating Latency
Symptom: When all prompts in a batch have different lengths, they get padded to the longest, and every request pays prefill cost for the longest prompt.
Root cause: Classic batching without PagedAttention must pad to equal length. PagedAttention solves this — but if you are measuring a non-vLLM engine, padding is a common cause of inflated latency.
Fix: State the padding strategy in the report, and preferably test with equal-length prompts (--random-input-len 1024).
4. Reporting Single-Batch Latency as Production Latency
Symptom: "Latency is 30 ms" — measured only at the batch=1 ideal case, while production at 32 concurrency runs at 200 ms.
Root cause: Single-batch latency is the engine's best case, not its production case. Production latency must be measured with concurrency.
Fix: Always report single-request latency and N-concurrency latency together; ideally plot a "concurrency vs latency" curve:
Concurrency p50 p99
1 22 25
4 24 30
16 28 50
32 35 180
64 50 1200 ← degradation begins hereThis is the "single-batch latency mistaken for production-grade latency" anti-pattern in Common Pitfalls and Anti-Patterns.
5. Not Stating Quantization Precision
Symptom: "vLLM runs Llama-2-7B at 1100 tok/s" — with an AWQ INT4 model, while the compared TGI runs bf16.
Root cause: Quantized and original models differ in throughput by 1.5–2× by nature — that is a model difference, not an engine difference.
Fix: Model version, quantization precision, and weight group size (q_group_size) must be stated explicitly, and both sides must use the same quantization precision in a comparison.
Self-check checklist
Before publishing any benchmark number, walk these five items:
- [ ] At least 3 warmup runs, then at least 100 measured runs
- [ ] Uses
torch.cuda.synchronize()or full HTTP round-trip time - [ ] Explicitly states prompt lengths and padding strategy
- [ ] Reports single-request latency plus latency at ≥ 3 concurrency levels
- [ ] Model weight hash, quantization precision, and CUDA version are all in the report
Miss one item and your numbers are not credible. See Common Pitfalls and Anti-Patterns and "end-to-end benchmarks, not single-operator benchmarks" in Deployment Design Principles.
5. Report Template
A qualified inference benchmark report should have five sections:
markdown
# Inference Benchmark Report: <model> on <hardware>
## 1. Environment
- Hardware: NVIDIA A100 80GB × 1
- Driver: 535.104.05, CUDA 12.4, cuDNN 9.0
- Engines: vLLM 0.6.3, TensorRT-LLM 0.13.0
- Model: Llama-2-7b-chat-hf (HuggingFace, sha256: ...)
- Quantization: bf16 (no quantization)
## 2. Configuration
- max_model_len: 4096
- max_num_seqs: 32
- gpu_memory_utilization: 0.9
- input_len: 1024, output_len: 256
- Concurrency levels: 1 / 4 / 16 / 32 / 64
- Warmup: 3 runs; test: 100 runs per level
## 3. Results
| Concurrency | Engine | p50 (ms) | p99 (ms) | Throughput (tok/s) | Peak memory (GB) |
|---|---|---|---|---|---|
| 1 | vLLM | 200 | 250 | 50 | 18 |
| 1 | TRT-LLM | 180 | 220 | 55 | 17 |
| 32 | vLLM | 35 | 180 | 1100 | 36 |
| 32 | TRT-LLM | 30 | 150 | 1300 | 32 |
## 4. Conclusions
- TRT-LLM is 10–20% faster than vLLM at every level
- But TRT-LLM has higher deployment complexity (engine build required)
- Online scenarios: vLLM recommended (development efficiency first);
offline scenarios: TRT-LLM recommended (throughput first)
## 5. Reproduction
- Code: `scripts/bench_vllm_vs_trtllm.sh`
- Model weights: HuggingFace `meta-llama/Llama-2-7b-chat-hf`
- Command: `bash scripts/bench_vllm_vs_trtllm.sh`Discipline of the template
All five sections are mandatory:
- Missing environment: no reproduction on a different machine
- Missing configuration: nobody knows how you tuned vLLM
- Missing results matrix: no basis for comparison decisions
- Missing conclusions: the report becomes a "data dump"
- Missing reproduction commands: the report is just "trust me"
See "how to read benchmark leaderboards" in Benchmark Data & Tool Profiles.
6. Turning Numbers into Decisions
A benchmark is not the destination; it is the input to a decision. Below are common patterns for turning numbers into decisions.
1. Engine Selection
After measuring the vLLM vs TRT-LLM comparison, the decision path:
Latency gap < 10% → pick vLLM (simpler deployment, larger community)
Latency gap 10–30% → depends on the team: NVIDIA background → TRT-LLM,
otherwise vLLM
Latency gap > 30% → latency-critical business → TRT-LLM
Throughput gap > 50% → offline batch processing → TRT-LLMSee Inference Engine Comparison.
2. Quantization Decision
After measuring AWQ INT4 vs bf16:
Throughput gain > 50%, MMLU drop < 1 → ship AWQ everywhere
Throughput gain 20–50%, MMLU drop 1–3 → gray-release at 20%, watch business metrics
Throughput gain < 20%, MMLU drop > 3 → skip AWQ, keep bf16See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
3. Concurrency Tuning
After measuring the "concurrency vs p99 latency" curve, find the inflection point:
1 → 22 ms
4 → 28 ms
16 → 35 ms ← before the inflection point
32 → 180 ms ← inflection point! p99 suddenly spikes
64 → 1200 ms ← severe degradationSet max_num_seqs = 16 (before the inflection point) and let excess requests queue or autoscale. See "max_num_seqs tuning" in Tuning and Performance Optimization.
4. Hardware Upgrade Decision
After measuring A100 vs H100:
Throughput gain < 1.5× → not worth upgrading (H100 costs 2×)
Throughput gain 1.5–2× → depends on whether business SLA is tight
Throughput gain > 2× → upgrade pays offSee Hardware Primer and GPU Architecture and Optimization.
7. Further Reading
- Latency, Throughput, and Concurrency — the theory behind this article's metrics
- Inference Engine Comparison — applying benchmark-driven decisions
- Deploy an Inference Service from Scratch — where this article's numbers come from
- Progressive Tutorial: Three Working Versions — one optimization per version, compared
- Tuning and Performance Optimization — turning benchmarks into tuning decisions
- Deployment Design Principles — end-to-end benchmark principles
- Common Pitfalls and Anti-Patterns — the full version of benchmark anti-patterns
- Benchmark Data & Tool Profiles — how to read public benchmark leaderboards
- Hardware Primer — cross-hardware comparison reference
- vLLM and PagedAttention — the origin of vLLM's bench tools
References
- vLLM: benchmarks/ — vLLM's bundled benchmark scripts
- MLCommons Inference: Language — the MLPerf-LLM standardized benchmark
- SGLang: bench — SGLang comparative benchmarks
- lm-evaluation-harness — model quality evaluation framework
- TensorRT-LLM: Benchmarks — TRT-LLM performance benchmarks
- OpenLLM Bench — comprehensive comparison leaderboard
- NVIDIA Nsight Systems — GPU-level profiling tool
- Prometheus + Grafana for LLM serving — production monitoring (complementary to benchmarking)