Skip to content

Benchmark Data & Tool Profiles

At a glance Benchmark data and tool profiles for inference acceleration — MLPerf Inference 4.1/5.0, public benchmarks of vLLM/SGLang/TensorRT-LLM, LLM throughput and latency across GPUs, INT8/INT4/FP8 comparisons, engine comparisons, and a quick reference of benchmarking tools.

This page contains time-sensitive content. Data is current as of 2026-08; engine versions, benchmark rankings, and product features may have changed since — verify against primary sources before citing.

Benchmark Data & Tool Profiles ​

Every conclusion in inference acceleration ultimately gets validated by a set of numbers: "how much faster with FP8," "how much faster is H200 than H100 at decode," "vLLM vs TensorRT-LLM — which wins." This page compiles publicly visible benchmark leaderboards and representative measurements as of 2026-08, paired with mainstream benchmarking tools. It is the "measured counterpart" of the Hardware Primer, the "data base" for Latency, Throughput, and Concurrency and Inference Benchmarking in Practice, and cross-validates Inference Engine Comparison.

Timeliness statement

All data on this page is current as of 2026-08, cited from official MLPerf results, engine READMEs/blogs, vendor white papers, and public papers. Benchmark numbers are strongly configuration-dependent (batch, sequence length, quantization, KV cache limits, engine version, hardware BIOS) — numbers from different sources are not directly comparable; always read them together with their configurations. After new releases ship, defer to the latest official results; this page serves only as a reference for "magnitudes and relative relationships."

How to use

This page has four blocks: leaderboards show where the SOTA sits; measured data tables show how fast a specific model is on a specific card; engine comparisons inform "which to pick"; tool quick reference shows "how to measure it yourself." Skim the leaderboards for the global picture, then look up measured data for your target model/hardware, and finally pick a tool and reproduce one run yourself.

1. Mainstream Benchmark Leaderboards ​

1.1 MLPerf Inference ​

MLPerf Inference is the industry's most authoritative inference benchmark, maintained by MLCommons in Datacenter and Edge tracks. LLMs became an official benchmark starting with Inference 4.0 (2024).

VersionRelease dateLLM test modelsKey scenarios
Inference 4.02024-04GPT-J 6B, Llama-2-70BSingle-stream and multi-stream latency, offline throughput
Inference 4.12024-09GPT-J 6B, Llama-2-70BAdded the Server scenario (sustained QPS under load)
Inference 5.02025-07Llama-2-70B, Llama-3.1-405B (FP8)405B enters the leaderboard for the first time; FP8 debuts at scale

Four scenarios:

  • SingleStream: single-request TTFT/TPOT — the latency view;
  • MultiStream: a small number of concurrent streams — latency + light throughput;
  • Server: maximum sustained QPS under a fixed SLA (e.g., P99 TTFT < 2s) — closest to online serving;
  • Offline: pure throughput in tokens/s under batch mode.

How to read MLPerf results

Compare only within the same version, scenario, and model — cross-version comparisons are invalid because benchmark rules change. NVIDIA typically submits the best-in-class results on H100/H200/B200 (FP8 + TensorRT-LLM), making it the "gold ruler" for hardware ceilings. Open-source engine results (vLLM/SGLang) mostly come from community submissions — close in magnitude, not necessarily optimal. See Inference Benchmarking in Practice.

1.2 Hugging Face Open LLM Leaderboard ​

The Hugging Face Open LLM Leaderboard v2 is the main leaderboard for model quality (not inference speed), covering reasoning, code, math, and tool use. It is the first reference for model selection — "is this model capable enough" is decided there; "how fast does it run" is decided by the other tables on this page. See Inference Benchmarking in Practice.

Don't confuse quality leaderboards with performance leaderboards

The Open LLM Leaderboard measures model quality (accuracy, pass@k); MLPerf measures inference performance (latency, throughput). A model can be high-quality but slow (405B FP16), or fast but low-quality (1B INT4). You must look at both.

1.3 Official engine benchmarks ​

EngineOfficial benchmark entryWhat it measures
vLLMbenchmark suiteonline serving throughput, offline inference, KV cache utilization, prefix caching
TensorRT-LLMBenchmark repomicrobenchmarks, end-to-end latency, throughput
SGLangsglang benchstructured generation, prefix cache hit rate, multi-round
TGIbenchmark scriptlatency, throughput, memory

Engine READMEs usually carry a "tokens/s for model X on card Y" number — the fastest entry point for reproducing a similar configuration. But configurations vary enormously: don't compare a vLLM blog number directly against a TensorRT-LLM blog number — batch, sequence length, quantization precision, and engine versions all differ. See Inference Engine Comparison and Inference Benchmarking in Practice.

2. Representative LLM Inference Data per GPU ​

The tables below aggregate representative data from public benchmarks. All numbers assume a single SXM card with configurations aligned as much as possible (same batch and sequence lengths). Differences of 10–30% across sources are normal — read magnitudes and relative relationships only, never as exact values.

2.1 Llama-3-8B inference throughput (tokens/s, single card) ​

GPUPrecisionDecode throughput (batch=1)Multi-stream throughput (batch=32)Data source
A100 80GBFP16~50 tok/s~1800 tok/svLLM blog
A100 80GBINT8 (W8A8)~70 tok/s~2600 tok/sTensorRT-LLM benchmark
H100 SXM 80GBFP16~110 tok/s~3800 tok/svLLM blog
H100 SXM 80GBFP8~160 tok/s~5500 tok/sNVIDIA LLM benchmark
H200 141GBFP8~220 tok/s~7800 tok/sNVIDIA MLPerf 5.0
B200FP4~400 tok/s~14000 tok/sNVIDIA Blackwell LLM paper
RTX 4090INT4 (AWQ)~80 tok/s~1200 tok/sllama.cpp community
MI300XFP8~180 tok/s~6200 tok/sAMD MI300X benchmark

Why is B200's number so high

B200's FP4 Tensor Core compute is ~9× H100's FP8 (sparse), and its HBM3e bandwidth of 8 TB/s is 2.4× H100's — while Llama-3-8B on B200 is fully bandwidth-limited (decode is memory-bound), so throughput gain ≈ bandwidth ratio × precision compression ratio. FP4 halves weight volume, halving per-token memory traffic — another doubling. See Hardware Primer and The Roofline Model and Compute Analysis.

2.2 Llama-3-70B inference throughput (tokens/s, single card) ​

A 70B model in FP16 needs ~140 GB — it doesn't fit on a single 80 GB card, so quantization or TP is mandatory. The table below shows single-card + INT4 quantization or TP=2 data.

GPUPrecisionDeploymentDecode throughput (batch=1)Multi-stream throughput (batch=32)
A100 80GBINT4 (AWQ)single card~18 tok/s~700 tok/s
H100 SXM 80GBINT4 (AWQ)single card~40 tok/s~1500 tok/s
H100 SXM 80GB ×2FP8TP=2~75 tok/s~2800 tok/s
H200 141GBINT4single card~55 tok/s~2100 tok/s
H200 141GBFP8single card~35 tok/s~1500 tok/s
B200FP4single card~120 tok/s~4500 tok/s
MI300XFP8single card~50 tok/s~2000 tok/s

Single-card vs dual-card 70B — the key difference

70B FP16 (140 GB) doesn't fit on a single 80 GB card — you must either quantize to INT4 (~35 GB) or use TP=2 (70 GB per card). Single-card INT4 has higher throughput but larger accuracy loss; dual-card FP8 has higher accuracy but requires NVLink. This is the core trade-off of 70B deployment — see Weight-Only Quantization and Mixed Precision and Distributed Inference (TP/PP).

2.3 Qwen2.5-72B latency across GPUs (TTFT / TPOT) ​

1024-token input prompt, 128-token output, single request.

GPUPrecisionTTFT (ms)TPOT (ms)E2E (s)
H100 SXM 80GBFP8 (single card, INT4 weights)280283.86
H100 SXM 80GB ×2FP8 (TP=2)180182.48
H200 141GBFP8 (single card)220223.04
B200FP4 (single card)120121.66
RTX 4090INT4 (AWQ)480486.62

2.4 INT8 vs INT4 vs FP8 (Llama-3-70B, H100 SXM) ​

PrecisionWeight sizeSingle-card memory (incl. KV cache)Decode throughput (batch=1)Quality loss
FP16 (baseline)140 GBdoesn't fit—0
FP8 (W8A8)70 GB~80 GB (just barely)~30 tok/s<1%
INT8 (W8A8 SmoothQuant)70 GB~80 GB~32 tok/s1–2%
INT4 (AWQ, weight-only)35 GB~45 GB~40 tok/s2–4%
INT4 (GPTQ)35 GB~45 GB~38 tok/s2–4%

How to pick a precision

H100/H200: prefer FP8 — native hardware support, smallest quality loss; A100/RTX 4090: pick INT4 (AWQ/GPTQ) — weight-only is the only option, but the memory and throughput advantages are largest; accuracy-sensitive scenarios: FP8, or keep key layers in FP16 (mixed precision). See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

3. Cross-Engine Comparison Benchmarks ​

Horizontal comparisons on the same hardware, same model, different engines best answer "which one should I pick." The table below aggregates representative community results for Llama-3-70B INT4 on H100 SXM 80GB.

EngineVersionPrecisionMulti-stream throughput (batch=32)TTFT (ms)Notes
vLLM0.6.xINT4 (AWQ)~1450 tok/s240PagedAttention + continuous batching
TensorRT-LLM0.13+INT4~1800 tok/s200in-flight batching + Planner
SGLang0.3.xINT4 (AWQ)~1500 tok/s230strong RadixAttention prefix caching
TGI2.xINT4 (AWQ)~1300 tok/s260same magnitude as vLLM

Common traps in engine comparisons

  1. Version differences: vLLM 0.5 vs 0.6 differ by 20%+ — must compare same versions;
  2. Different quantization implementations: both may say "INT4 AWQ," but TensorRT-LLM uses hardware dequantize GEMM while vLLM uses the Marlin kernel — a 10%-scale gap;
  3. Different batching policies: implementation details of continuous batching (e.g., whether chunked prefill is on) matter enormously;
  4. First-token vs multi-stream trade-off: an engine with low TTFT may have lower throughput (small prefill chunks), and vice versa. When measuring yourself, you must fix batch, sequence length, prompt distribution, and engine version. See Inference Engine Comparison and Inference Benchmarking in Practice.

3.1 Engine selection by model scale ​

Model scaleFits single cardRecommended engineKey optimizations
7B–8B FP1616 GB — single card sufficesvLLMcontinuous batching + PagedAttention
13B–14B INT4~8 GB, single cardvLLM / llama.cppINT4 on the edge
30B–34B INT4~18 GB, single cardvLLM / SGLangprefix caching
70B INT4~40 GB, single cardvLLM / TensorRT-LLMFP8 preferred (Hopper)
70B FP8~70 GB, single H200TensorRT-LLMin-flight batching
70B FP16 / 405B INT4multi-card TPTensorRT-LLM / vLLMNVLink + TP=2/4/8
405B FP88×H200 NVL / GB200TensorRT-LLMNVLink Switch full interconnect

4. Benchmarking Tool Quick Reference ​

4.1 Engine-built-in benchmarks ​

ToolEngineWhat it measuresEntry point
benchmark_throughput.pyvLLMoffline throughputvllm/benchmarks/
benchmark_serving.pyvLLMonline serving latency and throughputvllm/benchmarks/
benchmark_latency.pyvLLMsingle-stream latencyvllm/benchmarks/
cpp/test/perfTensorRT-LLMmicrobenchmarksTensorRT-LLM/benchmarks/
sglang.benchSGLangone-command multi-scenario stress testingsglang/benchmark/
text-generation-benchmarkTGITGI server-side stress testingtext-generation-inference/benchmark/

4.2 General-purpose benchmarking tools ​

ToolPurposeEntry point
MLPerf Inferencethe industry-standard benchmarkmlcommons.org
lm-evaluation-harnessmodel quality evaluation (used by the Open LLM Leaderboard)EleutherAI
vLLM benchmark suiteLLM serving stress testingvllm-project
SGLang benchstructured-generation stress testingsglang-project
promptbenchunified multi-model multi-task evaluationMicrosoft
opencompassShanghai AI Laboratory's comprehensive evaluation frameworkopen-compass

4.3 Profiling tools ​

ToolWhat it measuresUsage
Nsight SystemsGPU timeline, kernel time shares, CPU-GPU asyncnsys profile python script.py
Nsight Computeper-kernel compute/bandwidth/occupancy analysisncu --set full python script.py
PyTorch ProfilerPyTorch operator-level time and memorytorch.profiler.profile
Triton tutorialcustom-kernel performance tuningOpenAI Triton

Nsight Systems is the first tool for diagnosing inference bottlenecks

80% of inference performance problems are visible in one nsys profile pass: "CPU waiting on the GPU," "kernel-launch gaps too large," "one operator taking 60% of the time" — all obvious on the timeline. See GPU Architecture and Optimization and Tuning and Performance Optimization.

4.4 Minimal configuration for your own benchmark ​

A comparable LLM inference benchmark must fix at least the following variables (details in Inference Benchmarking in Practice):

VariableExample valueWhy it matters
HardwareH100 SXM 80GBthe hard ceiling of compute and bandwidth
ModelLlama-3-70Bparameter count determines memory traffic
PrecisionINT4 AWQaffects both memory and compute
Batch size1 / 32single-stream latency vs multi-stream throughput
Input length1024 tokensprefill compute
Output length128 tokensdecode repetitions
Engine versionvLLM 0.6.3same engine varies 20% across versions
KV cache limit90% of memorycaps the batch ceiling
Prefix cachingoffotherwise identical prompts hit the cache

5. Common Misreadings of Benchmark Data ​

Five traps behind the numbers

  1. Peak ≠ measured: 1979 TFLOPS FP8 is a sparse peak — hitting 30–50% in real LLM decode already makes you a top performer;
  2. Single card ≠ linear multi-card scaling: TP=2 does not double throughput — AllReduce overhead and load balancing cap the speedup at 1.6–1.8×;
  3. batch=1 and batch=32 are different worlds: the former is memory-bound, the latter compute-bound — conclusions can be opposite;
  4. Prompt distribution matters enormously: prompts of the same length but different token distributions can differ 2× in prefill time (long repeated token sequences hit the cache);
  5. Cold start ≠ steady state: post-warmup throughput typically rises 10–20% — benchmarks must warm up first.

Further Reading ​