Skip to content

Inference Benchmarking in Practice

At a glance How do you measure, report, and cheat-proof inference optimization numbers? This article presents a methodology for LLM inference benchmarking: what to measure (throughput / TTFT / TPOT / p99 / memory), what to measure with (vLLM offline / mlperf-llm / sglang bench / lm-eval-harness), how to control variables, and how to write the report — with an anti-cheat checklist and a report template.

Inference Benchmarking in Practice ​

A wrong number is more dangerous than no number at all — it dresses an optimization decision in "science" while hiding the real bottleneck.

There is an old saying in the inference optimization world: "my model runs faster than yours on your hardware" — that is marketing talk, not an engineering conclusion. The same model on the same GPU can produce tokens/s numbers that differ by 3× between two teams, and the reason is the test method: different batch sizes, different prompt lengths, different quantization precision, different KV cache configurations, no cold/warm start separation. A benchmark without methodology is worse than no benchmark — it keeps you investing in the wrong direction.

This article lays out a reusable, comparable, cheat-proof methodology for LLM inference benchmarking. After reading it, you should be able to write a test report that anyone can reproduce for any two-engine comparison, recognize the five most common forms of "benchmark cheating," and turn test results into decision criteria. This is the source methodology for the performance numbers in Deploy an Inference Service from Scratch and Progressive Tutorial: Three Working Versions.

Scope of this article

This article focuses on single-instance LLM inference benchmarking — engine comparisons on a single machine with one or more GPUs. Distributed system-level benchmarking (cluster QPS, cross-instance latency, network overhead) is a separate methodology and is not covered here. See Triton Inference Server and Model Serving and Orchestration.

1. What to Measure: Core Metrics for LLM Inference ​

The biggest difference between LLM inference and classic ML inference: the metric is not one number but a set. Looking only at tokens/s is like looking only at an average score — it hides long-tail latency, hides time to first token, and hides concurrency degradation. The five metrics below are all mandatory.

1. The Five Core Metrics ​

MetricDefinitionBusiness meaningWhat it measures
TTFT (Time To First Token)Time from request to first received token"Responsiveness" in chat scenarios; the boundary of user patiencePrefill phase + queue wait
TPOT (Time Per Output Token)Average per-token time during generation"Typing speed" of streaming generationDecode phase + KV cache access
End-to-end latency (E2E)Time from request to complete output"Wait-for-result" scenarios: batch jobs, document summarizationTTFT + N × TPOT
Throughput (tokens/s)Total tokens generated per unit timeCluster throughput capacity, cost accounting(Concurrency × output_len) / total time
Peak GPU memoryMaximum memory occupied during inferenceDeployment feasibility, concurrency ceilingWeights + KV pool + temporaries

2. Latency Distribution: p50 / p95 / p99 ​

Average latency lies — 28 of 30 concurrent requests take 1 s and 2 take 10 s, so the average of 1.6 s looks fine, but those two 10-second requests have already lost the user in a chat scenario. Production must look at percentiles:

Latency distribution (32 concurrency, sorted by TPOT):
p50  = 22 ms    ← median, "the normal-case experience"
p90  = 35 ms    ← 90% of requests are faster than this
p95  = 50 ms    ← the "long tail" SREs care about
p99  = 180 ms   ← the truly "unlucky users"
p999 = 1200 ms  ← extreme cases when the system is degrading
  • p50: representative of user experience, but hides the long tail
  • p95: the SRE's primary view — it represents "the floor most users get"
  • p99: watch this in cost-sensitive scenarios — pushing p99 from 200 ms to 180 ms may cost double the machines
  • p999: usually not reported, because the extreme 0.1% cases tend to come from GC, network jitter, and other non-model factors

Discipline on averages vs. percentiles

  • External reporting: report p50 + p99 (a pair reflecting center and tail)
  • Internal SRE: report p50/p90/p95/p99/p999 (a set reflecting the shape of the distribution)
  • Never report only the average — one slow request gets averaged away by 99 fast ones, and the problem disappears from view

See "latency distribution and queuing theory" in Latency, Throughput, and Concurrency.

3. The Three Layers of Peak Memory ​

Memory is not one number; it is three layers:

LayerContentsOptimized by
WeightsThe model parameters themselvesQuantization, pruning, distillation (see Model Quantization Fundamentals)
KV cacheK/V tensors of past tokensPagedAttention, block_size, max_num_seqs
TemporariesAttention intermediate matrices, activationsKernel fusion, FlashAttention

The "memory usage" you see in nvidia-smi is the sum of all three — without looking at each layer, you are tuning blind: "Why does a 14 GB model occupy 36 GB?" Because the KV pool takes 22 GB. See The GPU Memory Hierarchy and the Bandwidth Wall.

2. What to Measure With: A Tour of Tools ​

Tools for LLM inference benchmarking fall into four categories; each measures different metrics and serves different purposes.

1. Tool Matrix ​

ToolPrimary useWhat it measuresMaintainer
vLLM benchmarks/vLLM's own performance benchmarking and comparisonsThroughput, latency, memoryvLLM project
MLPerf-LLMStandardized benchmark across hardware and enginesThroughput, latency, costMLCommons
SGLang bench/SGLang vs. vLLM comparisonsThroughput, latency, complex schedulingSGLang project
lm-eval-harnessModel quality evaluation (not performance)MMLU/HellaSwag and other accuracy metricsEleutherAI
trtllm-benchTensorRT-LLM performance benchmarkingThroughput, latencyNVIDIA
openllm-benchCross-engine, cross-quantization comparisonAll dimensionsCommunity

2. vLLM Offline Inference Benchmark (Most Common) ​

vLLM's bundled benchmarks/benchmark_throughput.py and benchmark_latency.py are the most common "entry-level" benchmarks:

bash
# Offline throughput benchmark (no HTTP; direct Python API)
python benchmarks/benchmark_throughput.py \
    --model meta-llama/Llama-2-7b-chat-hf \
    --backend vllm \
    --input-len 1024 \
    --output-len 256 \
    --num-prompts 512 \
    --dtype bfloat16 \
    --max-model-len 4096

# Single-request latency benchmark
python benchmarks/benchmark_latency.py \
    --model meta-llama/Llama-2-7b-chat-hf \
    --backend vllm \
    --input-len 1024 \
    --output-len 256 \
    --batch-size 1 \
    --num-iters 100

The advantage: it talks directly to vLLM's internal API with no HTTP layer overhead, so it measures the engine itself. The downside: it only measures vLLM (other engines need their own scripts). See vLLM and PagedAttention.

3. MLPerf-LLM (Standardized Comparison) ​

MLPerf-LLM is a standardized benchmark across hardware and engines — all participants test under identical rules, and results are publicly comparable:

DimensionWhat it measures
Single-streamSingle-request latency (TTFT + TPOT)
Multi-streamLatency distribution under concurrent requests
OfflinePure throughput (no concurrency constraints)
ServerOnline QPS and SLO attainment

MLPerf's discipline is strict: every participant must report the full configuration (hardware model, driver version, engine version, model weight hash), and results must be reproducible. This is the gold standard for industrial-grade comparison — if your vLLM number of 1100 tok/s on an A100 can't be reconciled with the public MLPerf leaderboard, the problem is your method.

4. SGLang Bench (Comparative Testing) ​

SGLang's bench/ provides comparison scripts for vLLM and SGLang — it covers complex-scheduling scenarios (multi-turn dialogue, long context, structured output) more thoroughly than vLLM's bundled scripts:

bash
# SGLang vs vLLM comparison
python -m sglang.bench_serving \
    --backend vllm \
    --base-url http://localhost:8000 \
    --model meta-llama/Llama-2-7b-chat-hf \
    --dataset-name random \
    --random-input-len 1024 \
    --random-output-len 256 \
    --num-prompts 512

See vLLM and PagedAttention and "complex scheduling scenarios" in Inference Engine Comparison.

5. lm-eval-harness (Model Quality) ​

Performance is not everything — after quantization, pruning, or distillation, has the model gotten dumber? That is lm-evaluation-harness's job:

bash
# Run the MMLU 5-shot evaluation
lm_eval --model vllm \
    --model_args pretrained=meta-llama/Llama-2-7b-chat-hf \
    --tasks mmlu \
    --num_fewshot 5 \
    --batch_size 8

# Run the AWQ-quantized model
lm_eval --model vllm \
    --model_args pretrained=./Llama-2-7b-chat-hf-awq,quantization=awq \
    --tasks mmlu,hellaswag,arc_challenge \
    --num_fewshot 5

Performance numbers must always be reported together with quality numbers — reporting tokens/s without MMLU is like reporting cost without effect. See "quantization is a trade-off" in Model Quantization Fundamentals.

3. Key Variables: What Moves the Numbers Significantly ​

Each of the variables below can change the numbers by more than 2×. A report must explicitly state the value of every variable, or it cannot be reproduced.

1. Variable Matrix ​

VariableDefaultDirection of impactTypical effect
batch_size / max_num_seqs1 / 256Larger → throughput ↑ but p99 latency ↑5–10×
input_len (prompt length)Not fixedLonger → TTFT ↑, KV cache ↑Linear
output_len (generation length)Not fixedLonger → per-request latency ↑, TPOT unchangedLinear
Concurrency1Higher → throughput ↑ but p99 latency ↑3–30×
Quantization precisionbf16INT4 → memory ↓ and TPOT ↓1.3–2×
KV cache size (gpu_mem_util × block_size)0.9 × 16Larger → admits more concurrency2–5×
ModelNot fixedLarger → every metric degrades5–10×
GPU modelA100 40GBNewer is faster (A100→H100: 2×)1.5–3×
CUDA / driver versionLatestMismatch can cost 30%1.0–1.5×

2. Discipline for Fair Comparisons ​

When comparing two engines, you must hold constant:

Same model (same weight hash)
Same GPU (same model, same driver)
Same batch_size or concurrency
Same input_len and output_len
Same quantization precision
Same max_model_len and KV pool
Single variable: the engine itself

Miss one variable and the comparison is unfair. The classic counterexample: "vLLM has PagedAttention on by default, TGI doesn't" — run both with default settings and vLLM wins by default, but what wins is not the engine, it is the KV pool configuration. A fair comparison must tune both sides to their respective best configurations, not their defaults.

Comparing "default configurations" is a scam

Most "vLLM is 5× faster than TGI" comparison articles are default-configuration comparisons — PagedAttention enabled on one side, disabled on the other. The real comparison is "best-tuned vLLM vs best-tuned TGI" — that is the state users will actually tune the product into. See Inference Engine Comparison.

4. Cheat-Proofing: The 5 Most Common Tricks ​

The five forms of "cheating" in benchmark circles — most are not intentional, just mis-measurement that went unnoticed:

1. Not Separating Cold Start from Warm Start ​

Symptom: The first run is slow, later runs are fast, and the report only cites the later ones.

Root cause: CUDA kernel compilation, CUDA Graph capture, and KV cache warmup all happen in the first few inferences. A cold start being 3–10× slower is normal.

Fix:

python
# Standard practice: warmup N times (N ≥ 3), then benchmark
for _ in range(3):
    _ = model.generate(**inputs, max_new_tokens=8)   # warmup

# Formal test
t0 = time.perf_counter()
out = model.generate(...)
t1 = time.perf_counter()

State in the report: "3 warmup runs, then 100 measured runs, p50/p99 reported."

2. Ignoring CPU-GPU Overlap ​

Symptom: Using time.time() to measure end-to-end, which includes Python interpretation, HTTP serialization, and other CPU time.

Root cause: CUDA is asynchronous — when model.generate returns, the GPU is still running. Timing without torch.cuda.synchronize() measures "the time Python takes to launch the GPU."

Fix:

python
torch.cuda.synchronize()
t0 = time.perf_counter()
out = model.generate(...)
torch.cuda.synchronize()    # ← wait for the GPU to actually finish
t1 = time.perf_counter()

For HTTP service tests, use the time when the client has received the last chunk, not the time of request.send.

3. Padding Inflating Latency ​

Symptom: When all prompts in a batch have different lengths, they get padded to the longest, and every request pays prefill cost for the longest prompt.

Root cause: Classic batching without PagedAttention must pad to equal length. PagedAttention solves this — but if you are measuring a non-vLLM engine, padding is a common cause of inflated latency.

Fix: State the padding strategy in the report, and preferably test with equal-length prompts (--random-input-len 1024).

4. Reporting Single-Batch Latency as Production Latency ​

Symptom: "Latency is 30 ms" — measured only at the batch=1 ideal case, while production at 32 concurrency runs at 200 ms.

Root cause: Single-batch latency is the engine's best case, not its production case. Production latency must be measured with concurrency.

Fix: Always report single-request latency and N-concurrency latency together; ideally plot a "concurrency vs latency" curve:

Concurrency   p50    p99
1             22     25
4             24     30
16            28     50
32            35     180
64            50     1200  ← degradation begins here

This is the "single-batch latency mistaken for production-grade latency" anti-pattern in Common Pitfalls and Anti-Patterns.

5. Not Stating Quantization Precision ​

Symptom: "vLLM runs Llama-2-7B at 1100 tok/s" — with an AWQ INT4 model, while the compared TGI runs bf16.

Root cause: Quantized and original models differ in throughput by 1.5–2× by nature — that is a model difference, not an engine difference.

Fix: Model version, quantization precision, and weight group size (q_group_size) must be stated explicitly, and both sides must use the same quantization precision in a comparison.

Self-check checklist

Before publishing any benchmark number, walk these five items:

  • [ ] At least 3 warmup runs, then at least 100 measured runs
  • [ ] Uses torch.cuda.synchronize() or full HTTP round-trip time
  • [ ] Explicitly states prompt lengths and padding strategy
  • [ ] Reports single-request latency plus latency at ≥ 3 concurrency levels
  • [ ] Model weight hash, quantization precision, and CUDA version are all in the report

Miss one item and your numbers are not credible. See Common Pitfalls and Anti-Patterns and "end-to-end benchmarks, not single-operator benchmarks" in Deployment Design Principles.

5. Report Template ​

A qualified inference benchmark report should have five sections:

markdown
# Inference Benchmark Report: <model> on <hardware>

## 1. Environment
- Hardware: NVIDIA A100 80GB × 1
- Driver: 535.104.05, CUDA 12.4, cuDNN 9.0
- Engines: vLLM 0.6.3, TensorRT-LLM 0.13.0
- Model: Llama-2-7b-chat-hf (HuggingFace, sha256: ...)
- Quantization: bf16 (no quantization)

## 2. Configuration
- max_model_len: 4096
- max_num_seqs: 32
- gpu_memory_utilization: 0.9
- input_len: 1024, output_len: 256
- Concurrency levels: 1 / 4 / 16 / 32 / 64
- Warmup: 3 runs; test: 100 runs per level

## 3. Results

| Concurrency | Engine | p50 (ms) | p99 (ms) | Throughput (tok/s) | Peak memory (GB) |
|---|---|---|---|---|---|
| 1 | vLLM | 200 | 250 | 50 | 18 |
| 1 | TRT-LLM | 180 | 220 | 55 | 17 |
| 32 | vLLM | 35 | 180 | 1100 | 36 |
| 32 | TRT-LLM | 30 | 150 | 1300 | 32 |

## 4. Conclusions
- TRT-LLM is 10–20% faster than vLLM at every level
- But TRT-LLM has higher deployment complexity (engine build required)
- Online scenarios: vLLM recommended (development efficiency first);
  offline scenarios: TRT-LLM recommended (throughput first)

## 5. Reproduction
- Code: `scripts/bench_vllm_vs_trtllm.sh`
- Model weights: HuggingFace `meta-llama/Llama-2-7b-chat-hf`
- Command: `bash scripts/bench_vllm_vs_trtllm.sh`

Discipline of the template

All five sections are mandatory:

  • Missing environment: no reproduction on a different machine
  • Missing configuration: nobody knows how you tuned vLLM
  • Missing results matrix: no basis for comparison decisions
  • Missing conclusions: the report becomes a "data dump"
  • Missing reproduction commands: the report is just "trust me"

See "how to read benchmark leaderboards" in Benchmark Data & Tool Profiles.

6. Turning Numbers into Decisions ​

A benchmark is not the destination; it is the input to a decision. Below are common patterns for turning numbers into decisions.

1. Engine Selection ​

After measuring the vLLM vs TRT-LLM comparison, the decision path:

Latency gap < 10%      → pick vLLM (simpler deployment, larger community)
Latency gap 10–30%     → depends on the team: NVIDIA background → TRT-LLM,
                         otherwise vLLM
Latency gap > 30%      → latency-critical business → TRT-LLM
Throughput gap > 50%   → offline batch processing → TRT-LLM

See Inference Engine Comparison.

2. Quantization Decision ​

After measuring AWQ INT4 vs bf16:

Throughput gain > 50%, MMLU drop < 1     → ship AWQ everywhere
Throughput gain 20–50%, MMLU drop 1–3    → gray-release at 20%, watch business metrics
Throughput gain < 20%, MMLU drop > 3     → skip AWQ, keep bf16

See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

3. Concurrency Tuning ​

After measuring the "concurrency vs p99 latency" curve, find the inflection point:

1  → 22 ms
4  → 28 ms
16 → 35 ms    ← before the inflection point
32 → 180 ms   ← inflection point! p99 suddenly spikes
64 → 1200 ms  ← severe degradation

Set max_num_seqs = 16 (before the inflection point) and let excess requests queue or autoscale. See "max_num_seqs tuning" in Tuning and Performance Optimization.

4. Hardware Upgrade Decision ​

After measuring A100 vs H100:

Throughput gain < 1.5×    → not worth upgrading (H100 costs 2×)
Throughput gain 1.5–2×    → depends on whether business SLA is tight
Throughput gain > 2×      → upgrade pays off

See Hardware Primer and GPU Architecture and Optimization.

7. Further Reading ​

References ​