Appearance
Performance Optimization and Capacity Planning
One-sentence definition: performance optimization means "maximizing throughput within a given latency budget", and capacity planning means "working out how many machines you need to survive the target traffic" — both share the same metric language: P50/P99, QPS, TTFT/TPOT.
Industry insight: the most common mistake in performance work is single-point thinking — "it's slow, so optimize the model", "it's stuck, so add machines". In real systems the bottleneck usually hides in queuing, serialization, IO, or scheduling, not in the GPU. Industry experience says measure first, optimize bottom-up: load test to locate which layer (model / engine / service / system) is the bottleneck, change one layer at a time, and back every change with load-test data. Another uncomfortable truth: "elastic autoscaling" without capacity planning is flying blind — you have no idea how many replicas to scale to, and your HPA thresholds are pure gut feel.
1. The Performance Metric System
1. Latency distribution: P50/P99/P999
| Metric | Meaning | How to use it |
|---|---|---|
| P50 | Half of all requests land below this | The median user experience |
| P99 | 99% of requests land below this | The main anchor for your service quality (SLA) commitment |
| P999 | 99.9% of requests land below this | Long-tail troubleshooting (GC, cold starts, network jitter) |
Why P99 and not the average
Average latency gets dragged up by a few extremely slow requests, masking the fact that "most requests are actually fast"; P99 reflects the experience of the worst 1% of users. SLOs are usually written as "P99 < 100ms, met 99% of the time". Averages almost never appear in an SLA.
2. Throughput and concurrency
- QPS (Queries Per Second): requests processed per second — the basic unit of capacity planning;
- Concurrency: the number of requests in flight at the same time. The relationship:
concurrency = QPS × average latency(Little's Law, the core formula of capacity planning — see Section 4); - GPU utilization: distinguish compute utilization from bandwidth utilization — inference frequently shows low compute utilization while bandwidth is saturated. That is not "idle".
3. The LLM-specific latency breakdown
- TTFT (Time To First Token): from request to the first token (the "responsiveness" users perceive);
- TPOT (Time Per Output Token): average time to generate each token (the "typing speed" users perceive);
- Overall latency = TTFT + TPOT × output length. LLM optimization must treat these separately. See LLM Inference Optimization.
2. Locating Bottlenecks: First, Find Out What's Holding You Back
1. A four-quadrant classification
text
Compute bound Bandwidth bound
GPU compute utilization 90%+ Low compute utilization, but high
bandwidth / VRAM traffic
─────────────────────────────────────────────────────────────────
Compute time grows superlinearly Compute time grows linearly
with batch size with model size
Server-side bottleneck IO bottleneck
Heavy queuing, CPU pegged Network / disk / object storage
takes most of the latency budget
CPU flame graphs show nvidia-smi shows the GPU idle
serialization / scheduling overhead2. Three tools for pinpointing the problem
| Tool | What to look at |
|---|---|
nvidia-smi / nvidia-smi dmon | VRAM usage, SM utilization, temperature, power (confirm the GPU is actually working) |
perf / py-spy / flame graphs | What the CPU side is busy with (serialization? GIL? scheduling?) |
| Application-level instrumentation (OpenTelemetry) | Latency share per stage (model vs queue vs network — see Monitoring and Observability) |
Don't be fooled by nvidia-smi
The "GPU-Util" in nvidia-smi is SM utilization by default. With small inference batches it may read 5%, but that doesn't mean the card is broken or there's nothing left to optimize — the workload is memory-bound and bandwidth is the real bottleneck. Use ncu (Nsight Compute) to analyze bandwidth.
3. Four Layers of Optimization: Model → Engine → Service → System
text
Layer 4 System : connection pooling, IO multiplexing, caching, NIC tuning
Layer 3 Service : batching, result caching, concurrency model, request scheduling
Layer 2 Engine : graph optimization, TensorRT/TVM compilation, kernel specialization
Layer 1 Model : quantization, distillation, pruning, model architectureRule of thumb for ordering: bottom-up is easier, top-down is cheaper. Quantization at the model layer can deliver 2-4x gains with the smallest change, so do it first; system-layer tuning typically yields 10-30% with diminishing returns. This section focuses on the service and system layers (the model layer is covered in Quantization and Distillation, Pruning, and Low-Rank Factorization; the engine layer in Inference: From Forward Pass to Inference Engines).
1. Batching: the number-one lever in online inference
- Static batching: the caller accumulates N items before sending; simple, but latency is uncontrolled;
- Dynamic batching: the server merges requests arriving within a few-millisecond window into one batch and runs it on the GPU in a single pass — built into Triton, and it can boost small-model throughput 5-10x;
- Continuous batching: LLM-specific, scheduling at token granularity so the GPU never sits idle; see LLM Inference Optimization.
text
Dynamic batching (non-LLM):
t=0ms request A │
t=2ms request B │──► batch(A,B,C) in one forward pass ──► three responses
t=4ms request C │ (batch_timeout=10ms)2. Result caching and merging similar requests
- Exact caching: requests with an identical
(model, input)return the cached result directly. In recommendation scenarios hit rates can reach 30-50% (hot items, repeated queries); - Similar-request merging (dedup): treat requests whose embedding distance is below a threshold as similar and return the cached representative — big gains, but handle with care (it affects correctness; set a business-level tolerance).
3. Concurrency model: async is king
- Use async IO (asyncio) instead of "one thread per request + blocking calls";
- Route inference calls through server-side batching wherever possible; don't let a single request monopolize a model invocation;
- When multiple models share an instance, isolate them with per-model queues so one slow model can't drag everything down (see Multi-Model Serving with NVIDIA Triton).
4. The system layer: the easily overlooked 10-30%
- Connection pool reuse (HTTP keep-alive / gRPC channel reuse) — avoid setting up a connection per request;
- Serialization optimization: JSON → Protobuf, or message compression (gzip); significant gains on large payloads;
- Async logging and metrics: never let synchronous logging block the request path;
- NIC/kernel tuning: TCP BBR, buffer sizing. For cross-region deployments the network is often the biggest chunk of your P99.
4. Capacity Planning: From Estimation to Autoscaling
1. The core formula (Little's Law)
text
Concurrency (requests in flight) = QPS × average latency (seconds)
Example: target 1000 QPS at 50ms average latency → required concurrency = 1000 × 0.05 = 50How much concurrency a single instance can handle → back out the instance count:
text
Instances = target concurrency / measured per-instance concurrency × 1.3 (30% safety margin)2. The complete capacity planning workflow
text
(1) Load test: stress a single instance to plot its latency-throughput curve
(the max QPS at which P99 still meets the target)
(methodology: /practice/load-testing)
(2) Convert: target QPS / per-instance max QPS × 1.3 = instances needed
(3) Instrument: track live QPS, P99, and concurrency to validate the formula
(4) Autoscale: HPA on CPU or custom metrics; for GPU services prefer QPS or queue depth
(5) Reserve for peaks: size for peak QPS on big sale days (Double 11), not the averageThree rules for load testing
- Use a real traffic distribution (replay production traffic), not uniform synthetic traffic;
- Test against P99, not the average — P99 meeting target is what the SLA measures;
- Push until the point where "P99 starts to degrade" and record the QPS there — that's your capacity ceiling. Full methodology in Load Testing and Capacity Planning.
3. Elasticity strategies
| Scaling dimension | Metric | Recommendation |
|---|---|---|
| Horizontal scaling | Request queue depth / QPS / custom metrics | Bounded queues; scale up gradually, scale down quickly (avoid flapping) |
| GPU-specific | vGPU or full cards | For small models, partition cards with MIG/vGPU |
| Time-based planning | Cron-driven scaling preset from traffic curves | Scale out 1 hour ahead of big sale events |
4. Recommended optimization order
text
Step 1: Load test to locate the bottleneck layer (don't guess)
Step 2: Model layer (quantization) — biggest gains, smallest change
Step 3: Service layer (batching, caching, async)
Step 4: Engine layer (TensorRT compilation) — deep specialization
Step 5: System layer (connection pools, network, IO)
Step 6: Capacity planning + elastic scaling — let resources follow trafficAfter each step, load test again and quantify the gain; change one layer at a time, otherwise you won't be able to attribute the improvements.
Trade-offs
| Decision | Options | How to choose |
|---|---|---|
| Latency vs throughput | Small batches protect latency / large batches boost throughput | Online: protect P99 latency; offline: max out throughput |
| Optimize vs add machines | Tuning saves money but takes time / machines save time but cost money | Quantize and tune the service layer first; scale only after that |
| P50 vs P99 | Optimize the median / optimize the tail | Anchor the SLA on P99; chase the tail with dedicated investigations |
| Cache gains vs correctness | Caching is fast / freshness matters | Depends on the business's tolerance for stale results |
| Scale fast vs scale safe | Aggressive scale-in saves money / conservative avoids oscillation | Scale in slowly (prevent oscillation), scale out fast (prevent timeouts) |
One-line summary: performance is measured, not tweaked — set metrics first, load test to locate the bottleneck, optimize layer by layer, size capacity with Little's Law, and make every dollar land on a measurable gain.
Further Reading
- Load Testing and Capacity Planning — end-to-end practice, from load-test tools to capacity formulas
- LLM Inference Optimization — TTFT/TPOT and continuous batching
- Quantization — the biggest source of gains at the model layer
- Monitoring and Observability — the four golden signals: latency / traffic / errors / saturation
- Multi-Model Serving with NVIDIA Triton — dynamic batching in production