Appearance
Benchmark Data & Tool Profiles
Every conclusion in inference acceleration ultimately gets validated by a set of numbers: "how much faster with FP8," "how much faster is H200 than H100 at decode," "vLLM vs TensorRT-LLM — which wins." This page compiles publicly visible benchmark leaderboards and representative measurements as of 2026-08, paired with mainstream benchmarking tools. It is the "measured counterpart" of the Hardware Primer, the "data base" for Latency, Throughput, and Concurrency and Inference Benchmarking in Practice, and cross-validates Inference Engine Comparison.
Timeliness statement
All data on this page is current as of 2026-08, cited from official MLPerf results, engine READMEs/blogs, vendor white papers, and public papers. Benchmark numbers are strongly configuration-dependent (batch, sequence length, quantization, KV cache limits, engine version, hardware BIOS) — numbers from different sources are not directly comparable; always read them together with their configurations. After new releases ship, defer to the latest official results; this page serves only as a reference for "magnitudes and relative relationships."
How to use
This page has four blocks: leaderboards show where the SOTA sits; measured data tables show how fast a specific model is on a specific card; engine comparisons inform "which to pick"; tool quick reference shows "how to measure it yourself." Skim the leaderboards for the global picture, then look up measured data for your target model/hardware, and finally pick a tool and reproduce one run yourself.
1. Mainstream Benchmark Leaderboards
1.1 MLPerf Inference
MLPerf Inference is the industry's most authoritative inference benchmark, maintained by MLCommons in Datacenter and Edge tracks. LLMs became an official benchmark starting with Inference 4.0 (2024).
| Version | Release date | LLM test models | Key scenarios |
|---|---|---|---|
| Inference 4.0 | 2024-04 | GPT-J 6B, Llama-2-70B | Single-stream and multi-stream latency, offline throughput |
| Inference 4.1 | 2024-09 | GPT-J 6B, Llama-2-70B | Added the Server scenario (sustained QPS under load) |
| Inference 5.0 | 2025-07 | Llama-2-70B, Llama-3.1-405B (FP8) | 405B enters the leaderboard for the first time; FP8 debuts at scale |
Four scenarios:
- SingleStream: single-request TTFT/TPOT — the latency view;
- MultiStream: a small number of concurrent streams — latency + light throughput;
- Server: maximum sustained QPS under a fixed SLA (e.g., P99 TTFT < 2s) — closest to online serving;
- Offline: pure throughput in tokens/s under batch mode.
How to read MLPerf results
Compare only within the same version, scenario, and model — cross-version comparisons are invalid because benchmark rules change. NVIDIA typically submits the best-in-class results on H100/H200/B200 (FP8 + TensorRT-LLM), making it the "gold ruler" for hardware ceilings. Open-source engine results (vLLM/SGLang) mostly come from community submissions — close in magnitude, not necessarily optimal. See Inference Benchmarking in Practice.
1.2 Hugging Face Open LLM Leaderboard
The Hugging Face Open LLM Leaderboard v2 is the main leaderboard for model quality (not inference speed), covering reasoning, code, math, and tool use. It is the first reference for model selection — "is this model capable enough" is decided there; "how fast does it run" is decided by the other tables on this page. See Inference Benchmarking in Practice.
Don't confuse quality leaderboards with performance leaderboards
The Open LLM Leaderboard measures model quality (accuracy, pass@k); MLPerf measures inference performance (latency, throughput). A model can be high-quality but slow (405B FP16), or fast but low-quality (1B INT4). You must look at both.
1.3 Official engine benchmarks
| Engine | Official benchmark entry | What it measures |
|---|---|---|
| vLLM | benchmark suite | online serving throughput, offline inference, KV cache utilization, prefix caching |
| TensorRT-LLM | Benchmark repo | microbenchmarks, end-to-end latency, throughput |
| SGLang | sglang bench | structured generation, prefix cache hit rate, multi-round |
| TGI | benchmark script | latency, throughput, memory |
Engine READMEs usually carry a "tokens/s for model X on card Y" number — the fastest entry point for reproducing a similar configuration. But configurations vary enormously: don't compare a vLLM blog number directly against a TensorRT-LLM blog number — batch, sequence length, quantization precision, and engine versions all differ. See Inference Engine Comparison and Inference Benchmarking in Practice.
2. Representative LLM Inference Data per GPU
The tables below aggregate representative data from public benchmarks. All numbers assume a single SXM card with configurations aligned as much as possible (same batch and sequence lengths). Differences of 10–30% across sources are normal — read magnitudes and relative relationships only, never as exact values.
2.1 Llama-3-8B inference throughput (tokens/s, single card)
| GPU | Precision | Decode throughput (batch=1) | Multi-stream throughput (batch=32) | Data source |
|---|---|---|---|---|
| A100 80GB | FP16 | ~50 tok/s | ~1800 tok/s | vLLM blog |
| A100 80GB | INT8 (W8A8) | ~70 tok/s | ~2600 tok/s | TensorRT-LLM benchmark |
| H100 SXM 80GB | FP16 | ~110 tok/s | ~3800 tok/s | vLLM blog |
| H100 SXM 80GB | FP8 | ~160 tok/s | ~5500 tok/s | NVIDIA LLM benchmark |
| H200 141GB | FP8 | ~220 tok/s | ~7800 tok/s | NVIDIA MLPerf 5.0 |
| B200 | FP4 | ~400 tok/s | ~14000 tok/s | NVIDIA Blackwell LLM paper |
| RTX 4090 | INT4 (AWQ) | ~80 tok/s | ~1200 tok/s | llama.cpp community |
| MI300X | FP8 | ~180 tok/s | ~6200 tok/s | AMD MI300X benchmark |
Why is B200's number so high
B200's FP4 Tensor Core compute is ~9× H100's FP8 (sparse), and its HBM3e bandwidth of 8 TB/s is 2.4× H100's — while Llama-3-8B on B200 is fully bandwidth-limited (decode is memory-bound), so throughput gain ≈ bandwidth ratio × precision compression ratio. FP4 halves weight volume, halving per-token memory traffic — another doubling. See Hardware Primer and The Roofline Model and Compute Analysis.
2.2 Llama-3-70B inference throughput (tokens/s, single card)
A 70B model in FP16 needs ~140 GB — it doesn't fit on a single 80 GB card, so quantization or TP is mandatory. The table below shows single-card + INT4 quantization or TP=2 data.
| GPU | Precision | Deployment | Decode throughput (batch=1) | Multi-stream throughput (batch=32) |
|---|---|---|---|---|
| A100 80GB | INT4 (AWQ) | single card | ~18 tok/s | ~700 tok/s |
| H100 SXM 80GB | INT4 (AWQ) | single card | ~40 tok/s | ~1500 tok/s |
| H100 SXM 80GB ×2 | FP8 | TP=2 | ~75 tok/s | ~2800 tok/s |
| H200 141GB | INT4 | single card | ~55 tok/s | ~2100 tok/s |
| H200 141GB | FP8 | single card | ~35 tok/s | ~1500 tok/s |
| B200 | FP4 | single card | ~120 tok/s | ~4500 tok/s |
| MI300X | FP8 | single card | ~50 tok/s | ~2000 tok/s |
Single-card vs dual-card 70B — the key difference
70B FP16 (140 GB) doesn't fit on a single 80 GB card — you must either quantize to INT4 (~35 GB) or use TP=2 (70 GB per card). Single-card INT4 has higher throughput but larger accuracy loss; dual-card FP8 has higher accuracy but requires NVLink. This is the core trade-off of 70B deployment — see Weight-Only Quantization and Mixed Precision and Distributed Inference (TP/PP).
2.3 Qwen2.5-72B latency across GPUs (TTFT / TPOT)
1024-token input prompt, 128-token output, single request.
| GPU | Precision | TTFT (ms) | TPOT (ms) | E2E (s) |
|---|---|---|---|---|
| H100 SXM 80GB | FP8 (single card, INT4 weights) | 280 | 28 | 3.86 |
| H100 SXM 80GB ×2 | FP8 (TP=2) | 180 | 18 | 2.48 |
| H200 141GB | FP8 (single card) | 220 | 22 | 3.04 |
| B200 | FP4 (single card) | 120 | 12 | 1.66 |
| RTX 4090 | INT4 (AWQ) | 480 | 48 | 6.62 |
2.4 INT8 vs INT4 vs FP8 (Llama-3-70B, H100 SXM)
| Precision | Weight size | Single-card memory (incl. KV cache) | Decode throughput (batch=1) | Quality loss |
|---|---|---|---|---|
| FP16 (baseline) | 140 GB | doesn't fit | — | 0 |
| FP8 (W8A8) | 70 GB | ~80 GB (just barely) | ~30 tok/s | <1% |
| INT8 (W8A8 SmoothQuant) | 70 GB | ~80 GB | ~32 tok/s | 1–2% |
| INT4 (AWQ, weight-only) | 35 GB | ~45 GB | ~40 tok/s | 2–4% |
| INT4 (GPTQ) | 35 GB | ~45 GB | ~38 tok/s | 2–4% |
How to pick a precision
H100/H200: prefer FP8 — native hardware support, smallest quality loss; A100/RTX 4090: pick INT4 (AWQ/GPTQ) — weight-only is the only option, but the memory and throughput advantages are largest; accuracy-sensitive scenarios: FP8, or keep key layers in FP16 (mixed precision). See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
3. Cross-Engine Comparison Benchmarks
Horizontal comparisons on the same hardware, same model, different engines best answer "which one should I pick." The table below aggregates representative community results for Llama-3-70B INT4 on H100 SXM 80GB.
| Engine | Version | Precision | Multi-stream throughput (batch=32) | TTFT (ms) | Notes |
|---|---|---|---|---|---|
| vLLM | 0.6.x | INT4 (AWQ) | ~1450 tok/s | 240 | PagedAttention + continuous batching |
| TensorRT-LLM | 0.13+ | INT4 | ~1800 tok/s | 200 | in-flight batching + Planner |
| SGLang | 0.3.x | INT4 (AWQ) | ~1500 tok/s | 230 | strong RadixAttention prefix caching |
| TGI | 2.x | INT4 (AWQ) | ~1300 tok/s | 260 | same magnitude as vLLM |
Common traps in engine comparisons
- Version differences: vLLM 0.5 vs 0.6 differ by 20%+ — must compare same versions;
- Different quantization implementations: both may say "INT4 AWQ," but TensorRT-LLM uses hardware dequantize GEMM while vLLM uses the Marlin kernel — a 10%-scale gap;
- Different batching policies: implementation details of continuous batching (e.g., whether chunked prefill is on) matter enormously;
- First-token vs multi-stream trade-off: an engine with low TTFT may have lower throughput (small prefill chunks), and vice versa. When measuring yourself, you must fix batch, sequence length, prompt distribution, and engine version. See Inference Engine Comparison and Inference Benchmarking in Practice.
3.1 Engine selection by model scale
| Model scale | Fits single card | Recommended engine | Key optimizations |
|---|---|---|---|
| 7B–8B FP16 | 16 GB — single card suffices | vLLM | continuous batching + PagedAttention |
| 13B–14B INT4 | ~8 GB, single card | vLLM / llama.cpp | INT4 on the edge |
| 30B–34B INT4 | ~18 GB, single card | vLLM / SGLang | prefix caching |
| 70B INT4 | ~40 GB, single card | vLLM / TensorRT-LLM | FP8 preferred (Hopper) |
| 70B FP8 | ~70 GB, single H200 | TensorRT-LLM | in-flight batching |
| 70B FP16 / 405B INT4 | multi-card TP | TensorRT-LLM / vLLM | NVLink + TP=2/4/8 |
| 405B FP8 | 8×H200 NVL / GB200 | TensorRT-LLM | NVLink Switch full interconnect |
4. Benchmarking Tool Quick Reference
4.1 Engine-built-in benchmarks
| Tool | Engine | What it measures | Entry point |
|---|---|---|---|
benchmark_throughput.py | vLLM | offline throughput | vllm/benchmarks/ |
benchmark_serving.py | vLLM | online serving latency and throughput | vllm/benchmarks/ |
benchmark_latency.py | vLLM | single-stream latency | vllm/benchmarks/ |
cpp/test/perf | TensorRT-LLM | microbenchmarks | TensorRT-LLM/benchmarks/ |
sglang.bench | SGLang | one-command multi-scenario stress testing | sglang/benchmark/ |
text-generation-benchmark | TGI | TGI server-side stress testing | text-generation-inference/benchmark/ |
4.2 General-purpose benchmarking tools
| Tool | Purpose | Entry point |
|---|---|---|
| MLPerf Inference | the industry-standard benchmark | mlcommons.org |
| lm-evaluation-harness | model quality evaluation (used by the Open LLM Leaderboard) | EleutherAI |
| vLLM benchmark suite | LLM serving stress testing | vllm-project |
| SGLang bench | structured-generation stress testing | sglang-project |
| promptbench | unified multi-model multi-task evaluation | Microsoft |
| opencompass | Shanghai AI Laboratory's comprehensive evaluation framework | open-compass |
4.3 Profiling tools
| Tool | What it measures | Usage |
|---|---|---|
| Nsight Systems | GPU timeline, kernel time shares, CPU-GPU async | nsys profile python script.py |
| Nsight Compute | per-kernel compute/bandwidth/occupancy analysis | ncu --set full python script.py |
| PyTorch Profiler | PyTorch operator-level time and memory | torch.profiler.profile |
| Triton tutorial | custom-kernel performance tuning | OpenAI Triton |
Nsight Systems is the first tool for diagnosing inference bottlenecks
80% of inference performance problems are visible in one nsys profile pass: "CPU waiting on the GPU," "kernel-launch gaps too large," "one operator taking 60% of the time" — all obvious on the timeline. See GPU Architecture and Optimization and Tuning and Performance Optimization.
4.4 Minimal configuration for your own benchmark
A comparable LLM inference benchmark must fix at least the following variables (details in Inference Benchmarking in Practice):
| Variable | Example value | Why it matters |
|---|---|---|
| Hardware | H100 SXM 80GB | the hard ceiling of compute and bandwidth |
| Model | Llama-3-70B | parameter count determines memory traffic |
| Precision | INT4 AWQ | affects both memory and compute |
| Batch size | 1 / 32 | single-stream latency vs multi-stream throughput |
| Input length | 1024 tokens | prefill compute |
| Output length | 128 tokens | decode repetitions |
| Engine version | vLLM 0.6.3 | same engine varies 20% across versions |
| KV cache limit | 90% of memory | caps the batch ceiling |
| Prefix caching | off | otherwise identical prompts hit the cache |
5. Common Misreadings of Benchmark Data
Five traps behind the numbers
- Peak ≠ measured: 1979 TFLOPS FP8 is a sparse peak — hitting 30–50% in real LLM decode already makes you a top performer;
- Single card ≠ linear multi-card scaling: TP=2 does not double throughput — AllReduce overhead and load balancing cap the speedup at 1.6–1.8×;
- batch=1 and batch=32 are different worlds: the former is memory-bound, the latter compute-bound — conclusions can be opposite;
- Prompt distribution matters enormously: prompts of the same length but different token distributions can differ 2× in prefill time (long repeated token sequences hit the cache);
- Cold start ≠ steady state: post-warmup throughput typically rises 10–20% — benchmarks must warm up first.
Further Reading
- Latency, Throughput, and Concurrency — definitions of benchmark metrics
- Inference Benchmarking in Practice — how to run benchmarks yourself
- Inference Engine Comparison — selecting among vLLM/TensorRT-LLM/SGLang/TGI
- Hardware Primer — hardware parameter reference
- Weight-Only Quantization and Mixed Precision — INT4/INT8/FP8 quantization details
- Model Quantization Fundamentals — quantization method principles
- The Roofline Model and Compute Analysis — unified compute and bandwidth analysis
- Glossary — definitions of the terms used here
- Tuning and Performance Optimization — optimization after benchmarking
- Common Pitfalls and Anti-Patterns — engineering lessons behind benchmarks