Appearance
Latency, Throughput, and Concurrency
Concept Definition: The "Impossible Triangle" of Inference Systems
An inference system has three classes of user-facing metrics: latency — how long a single request must wait; throughput — how much work gets done per unit of time; and concurrency — how many users can be served simultaneously. There is a built-in trade-off among the three: blindly optimizing one typically degrades another. This tension is the central problem that Batching and Request Scheduling and Model Serving and Orchestration exist to resolve.
Three key insights for understanding these metrics:
- Latency and throughput are not the same thing — a system can be "fast per request but low in total throughput" (small batch), or "high in total throughput but slow per request" (large batch);
- Concurrency is the bridge between them — by Little's Law, concurrency = throughput × latency: know any two and you can derive the third;
- LLM inference has two-phase latency — prefill (compute-intensive) and decode (memory-intensive) have completely different bottlenecks and must be analyzed separately.
Inference optimization without a solid grasp of latency and throughput goes wrong in one of two ways: "a great QPS number that still feels laggy to users," or "low latency with GPU utilization under 20%" — both classic ways to fail in practice.
1. The Three Latency Time Points
Unlike a traditional classification model that returns output in one shot, LLM inference generates tokens in a stream. So latency must be split into three time points:
| Metric | Full Name | Meaning | User Perception |
|---|---|---|---|
| TTFT | Time To First Token | From request arrival to the first output token | "How fast it responds" — most critical for interactive scenarios |
| TPOT | Time Per Output Token | Average generation time for each subsequent token | "How fast it types" |
| E2E Latency | End-to-End Latency | Total time from request to completed output | What batch-processing scenarios care about |
The relationship among the three:
text
E2E = TTFT + TPOT × (output_tokens - 1)Why TTFT Is Measured Separately
TTFT corresponds to the prefill phase — the model processes the entire prompt (hundreds to thousands of tokens) in parallel, which is compute-intensive with a large one-off cost. Every subsequent token goes through the decode phase — autoregressive generation of one token, which is memory-intensive with a small per-step cost that repeats N times. The two phases have fundamentally different bottleneck profiles (see The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis), so latency must be measured separately.
Interactive vs. Batch: Choosing What to Optimize
Different scenarios have different sensitivities to the three metrics:
| Scenario | Key Metric | Compromisable Metric | Typical Requirement |
|---|---|---|---|
| Chat / customer service | TTFT | E2E (users tolerate a few seconds) | TTFT < 500ms, TPOT < 50ms |
| Code completion | TPOT | Throughput | TPOT < 30ms (keep up with typing) |
| Batch scoring | Throughput | Per-request latency | Several seconds of E2E per item is acceptable |
| Offline bulk generation | Throughput | Latency | Single-request latency doesn't matter at all |
Ask "How Do Users Actually Use It" Before Picking Metrics
Latency optimization for interactive scenarios (small batch, continuous batching) and throughput optimization for batch scenarios (large batch, static batching) are opposite strategies. Tuning against the wrong metric is wasted work. See Batching and Request Scheduling.
2. The Many Ways to Measure Throughput
Throughput is "amount processed per unit time," but "amount" has multiple definitions that must be kept straight:
| Measure | Formula | Applies To | Pitfall |
|---|---|---|---|
| QPS | Requests / second | Short requests, fixed output length | Mixing long and short requests distorts it |
| tokens/s (output) | Total output tokens / second | The most common measure in LLM inference | Output-only or input+output? The convention must be stated |
| tokens/s (input + output) | (Input + output) / second | Reflects total compute consumption | Disconnected from "how fast it feels" to users |
| concurrent users | Simultaneously online users | Capacity planning | Users ≠ requests (users may be idle) |
The de facto standard in the LLM world is output tokens/s — it directly answers "how much useful content can this hardware produce per second." But even then, you must also report input token count and average output length, or you can be fooled by the "inflate throughput with long outputs" trick.
3. Little's Law: The Universal Formula for the Three Metrics
Little's Law (John Little, 1961) from queueing theory gives the cleanest relationship among the three:
text
L = λ × W
i.e., concurrency = throughput × average latency- L: the average number of requests in the system at once (concurrency);
- λ: arrival rate (throughput, requests/second);
- W: the average time a request spends in the system (latency, seconds).
Example: your service handles 100 requests per second (throughput 100 QPS), and each request spends an average of 2 seconds in the system (latency 2s). Then on average there are 100 × 2 = 200 concurrent requests in the system. Know any two, derive the third — measure any two and the remaining one follows.
Why Little's Law Matters
- Capacity planning: if one GPU sustains 200 concurrent requests at an average latency of 2s, you can derive its throughput as 100 QPS, and compute how many GPUs you need to carry 1000 QPS;
- Diagnosing bottlenecks: if measured throughput is far below "theoretical throughput" while latency looks normal → the system is stalling on synchronization or queueing; if latency spikes while throughput barely grows → you've hit the concurrency ceiling (the GPU is full and requests are queueing);
- It holds for any stable system — no matter how complex your scheduling policy, this relationship cannot be escaped.
4. The Latency-Throughput Curve: The Hardware Baseline
Every inference platform has a latency-throughput curve — it describes "at a given concurrency level, what is the per-request latency." The typical shape:
text
Latency
│ ╱────── Saturation region
│ ╱
│ ╱
│ ╱
│ ╱
│ ╱
│ ╱
│ ╱──── Linear region
│ ╱
│ ╱
│ ╱──── Low-concurrency region
└────────────────────────────────────────→ Throughput / ConcurrencyThe curve has three segments:
- Low-concurrency region (bottom-left): the GPU is underfed — throughput grows linearly with concurrency while latency barely changes — adding concurrency is free;
- Linear region (middle): throughput keeps growing linearly with concurrency, but latency starts to rise (queueing + larger batches);
- Saturation region (top-right): the GPU is full — throughput stops growing, and every additional request turns into queueing delay — latency explodes.
Don't Operate in the Saturation Region
Many teams "added GPUs and it's still slow" — the root cause is that they're already in the saturation region, where extra concurrency just becomes latency. The right move is to find the knee of the curve (the concurrency level where throughput is near peak and latency starts climbing steeply) and keep production load around that point, instead of blindly piling on concurrency.
Why the LLM Curve Is Steeper
Traditional DNN inference curves are fairly flat — from batch 1 to 32, latency may only rise 30%. The LLM inference curve is much steeper, for two reasons:
- Prefill phase: the longer the prompt, the larger the per-step computation; batching multiple long prompts together blows through GPU memory and compute;
- KV cache memory: every concurrent request must hold its own KV cache (see The GPU Memory Hierarchy and the Bandwidth Wall) — the more concurrency, the greater the memory pressure, and once swapping kicks in, latency explodes several-fold.
So the core job of an LLM inference scheduler (like vLLM and PagedAttention) is to maximize concurrency within the memory budget without triggering swap.
5. Prefill vs. Decode: Two-Phase Latency Characteristics
LLM inference latency is special because it has two completely distinct phases:
| Phase | Compute/Memory Character | Main Bottleneck | Latency vs. Batch Size |
|---|---|---|---|
| Prefill (process the prompt) | Compute-intensive (large GEMM) | GPU compute (Tensor Core) | Grows almost linearly (compute is amortized) |
| Decode (generate tokens) | Memory-intensive (every generated token must load all weights) | HBM bandwidth | Nearly flat (bandwidth is sufficient; a larger batch only reads more KV cache) |
This is why LLM inference systems often use prefill/decode disaggregated scheduling — prefill batched separately to saturate GPU compute; decode merged into large batches to saturate bandwidth. See Batching and Request Scheduling.
text
Latency
│ prefill (compute-bound, grows linearly with batch)
│ ╱
│ ╱
│ ╱ decode (memory-bound, barely grows with batch)
│ ╱ ──────────────────────────
│ ╱
│ ╱
└────────────────────────────→ batch sizeThe Engineering Meaning of "TPOT Barely Changes with Batch"
It means piling on concurrency during decode is almost free — going from batch 1 to 32 may increase TPOT by only 10%, while throughput gains 30×. This is the foundation of the vLLM and PagedAttention continuous batching design: pack as many requests as possible into decode and squeeze the bandwidth dry.
6. Setting Concurrency: Theory vs. Practice
The theoretically optimal concurrency is "the concurrency at the knee of the curve." In practice, consider:
- Memory ceiling: every concurrent request occupies KV cache — the estimation formula is in The GPU Memory Hierarchy and the Bandwidth Wall. Without enough memory, no level of concurrency will run;
- Request length distribution: when long and short requests mix, long requests hold concurrency slots hostage (the core motivation for Batching and Request Scheduling's continuous batching);
- Queueing delay: theoretical latency counts only service time; in reality there is also queueing delay — higher concurrency makes queues longer, which in turn inflates E2E latency;
- Tail latency: p50 may look fine while p99 has already exploded — production systems must watch percentiles, not just averages.
7. Measurement: Which Tools to Use
"What's my model's latency and throughput?" — this must be answered with standardized benchmark tools, not by feel:
| Tool | Purpose | Characteristics |
|---|---|---|
| vLLM benchmark | Real-world LLM serving measurement | Real HTTP calls; full TTFT/TPOT/E2E/throughput metrics |
| MLPerf Inference LLM | Standardized cross-vendor comparison | Strict rules, fixed datasets — highly comparable but a high bar |
| Custom benchmark | End-to-end business scenarios | Must simulate the real request distribution (prompt length, QPS, concurrency) |
| PyTorch Profiler / Nsight | Operator-level bottleneck diagnosis | Not a user-facing metric — an optimizer's tool |
Three Big Benchmark Pitfalls
- The request distribution must be realistic: throughput measured with average prompt length 256 can be off by 5× from your production traffic that averages 1200;
- Warm up before measuring: the first requests include model loading and CUDA kernel compilation, which inflate latency — always run a few dozen warmup requests before collecting statistics;
- Look at percentiles: average latency alone is masked by occasional long tails; production systems watch p95/p99, and critical businesses watch p99.9.
More measurement practice in Inference Benchmarking in Practice and Common Pitfalls and Anti-Patterns.
8. Trade-offs
- Latency vs. throughput: interactive scenarios prioritize latency (small batch + continuous batching); batch scenarios prioritize throughput (large batch + static/dynamic batching);
- Concurrency vs. memory: higher concurrency means better throughput, but KV cache memory grows linearly — either add memory (H200 141GB), use Model Quantization Fundamentals to shrink weights and free up space, or use PagedAttention (vLLM and PagedAttention) to reduce fragmentation;
- TTFT vs. TPOT: many optimizations (such as speculative sampling) trade a bit of TTFT for a large drop in TPOT — interactive scenarios are sensitive to TTFT, so weigh carefully;
- Precision vs. speed: Weight-Only Quantization and Mixed Precision (W4A16) can double throughput at the cost of a small accuracy loss — worth it for most business scenarios.
Further Reading
- Batching and Request Scheduling — the engineering method for tuning concurrency to optimum
- The GPU Memory Hierarchy and the Bandwidth Wall — why piling on concurrency in decode is almost free
- The Roofline Model and Compute Analysis — determining whether an operator is compute- or memory-bound
- Model Serving and Orchestration — wiring latency/throughput metrics into a monitoring loop
- vLLM and PagedAttention — the industrial implementation of continuous batching
- Inference Benchmarking in Practice — how to build a trustworthy measurement pipeline
- Common Pitfalls and Anti-Patterns — common ways latency/throughput measurement goes wrong
References
- Little. A Proof for the Queuing Formula L = λW (Operations Research, 1961) — the original Little's Law paper
- MLPerf Inference: Benchmark Suite for ML Inference — the industry-standard inference benchmark
- vLLM: PagedAttention official documentation and benchmark scripts — industrial-grade LLM serving and measurement
- Gunho et al. Attention Is All You Need (NeurIPS 2017) — the origin of the two-phase prefill/decode behavior
- Pope et al. Efficiently Scaling Transformer Inference (MLSys 2023) — a systematic analysis of LLM inference latency-throughput curves