Skip to content

Performance Optimization and Capacity Planning

At a glance Latency, throughput, and cost form a three-way trade-off in inference serving. This article lays out a performance metric system (P50/P99/QPS/TTFT/TPOT), four layers of optimization (model → engine → service → system), and a capacity planning methodology.

Performance Optimization and Capacity Planning ​

One-sentence definition: performance optimization means "maximizing throughput within a given latency budget", and capacity planning means "working out how many machines you need to survive the target traffic" — both share the same metric language: P50/P99, QPS, TTFT/TPOT.

Industry insight: the most common mistake in performance work is single-point thinking — "it's slow, so optimize the model", "it's stuck, so add machines". In real systems the bottleneck usually hides in queuing, serialization, IO, or scheduling, not in the GPU. Industry experience says measure first, optimize bottom-up: load test to locate which layer (model / engine / service / system) is the bottleneck, change one layer at a time, and back every change with load-test data. Another uncomfortable truth: "elastic autoscaling" without capacity planning is flying blind — you have no idea how many replicas to scale to, and your HPA thresholds are pure gut feel.

1. The Performance Metric System ​

1. Latency distribution: P50/P99/P999 ​

MetricMeaningHow to use it
P50Half of all requests land below thisThe median user experience
P9999% of requests land below thisThe main anchor for your service quality (SLA) commitment
P99999.9% of requests land below thisLong-tail troubleshooting (GC, cold starts, network jitter)

Why P99 and not the average

Average latency gets dragged up by a few extremely slow requests, masking the fact that "most requests are actually fast"; P99 reflects the experience of the worst 1% of users. SLOs are usually written as "P99 < 100ms, met 99% of the time". Averages almost never appear in an SLA.

2. Throughput and concurrency ​

  • QPS (Queries Per Second): requests processed per second — the basic unit of capacity planning;
  • Concurrency: the number of requests in flight at the same time. The relationship: concurrency = QPS × average latency (Little's Law, the core formula of capacity planning — see Section 4);
  • GPU utilization: distinguish compute utilization from bandwidth utilization — inference frequently shows low compute utilization while bandwidth is saturated. That is not "idle".

3. The LLM-specific latency breakdown ​

  • TTFT (Time To First Token): from request to the first token (the "responsiveness" users perceive);
  • TPOT (Time Per Output Token): average time to generate each token (the "typing speed" users perceive);
  • Overall latency = TTFT + TPOT × output length. LLM optimization must treat these separately. See LLM Inference Optimization.

2. Locating Bottlenecks: First, Find Out What's Holding You Back ​

1. A four-quadrant classification ​

text
        Compute bound                        Bandwidth bound
   GPU compute utilization 90%+         Low compute utilization, but high
                                        bandwidth / VRAM traffic
   ─────────────────────────────────────────────────────────────────
   Compute time grows superlinearly     Compute time grows linearly
   with batch size                      with model size

        Server-side bottleneck               IO bottleneck
   Heavy queuing, CPU pegged            Network / disk / object storage
                                        takes most of the latency budget
   CPU flame graphs show                nvidia-smi shows the GPU idle
   serialization / scheduling overhead

2. Three tools for pinpointing the problem ​

ToolWhat to look at
nvidia-smi / nvidia-smi dmonVRAM usage, SM utilization, temperature, power (confirm the GPU is actually working)
perf / py-spy / flame graphsWhat the CPU side is busy with (serialization? GIL? scheduling?)
Application-level instrumentation (OpenTelemetry)Latency share per stage (model vs queue vs network — see Monitoring and Observability)

Don't be fooled by nvidia-smi

The "GPU-Util" in nvidia-smi is SM utilization by default. With small inference batches it may read 5%, but that doesn't mean the card is broken or there's nothing left to optimize — the workload is memory-bound and bandwidth is the real bottleneck. Use ncu (Nsight Compute) to analyze bandwidth.

3. Four Layers of Optimization: Model → Engine → Service → System ​

text
Layer 4  System  : connection pooling, IO multiplexing, caching, NIC tuning
Layer 3  Service : batching, result caching, concurrency model, request scheduling
Layer 2  Engine  : graph optimization, TensorRT/TVM compilation, kernel specialization
Layer 1  Model   : quantization, distillation, pruning, model architecture

Rule of thumb for ordering: bottom-up is easier, top-down is cheaper. Quantization at the model layer can deliver 2-4x gains with the smallest change, so do it first; system-layer tuning typically yields 10-30% with diminishing returns. This section focuses on the service and system layers (the model layer is covered in Quantization and Distillation, Pruning, and Low-Rank Factorization; the engine layer in Inference: From Forward Pass to Inference Engines).

1. Batching: the number-one lever in online inference ​

  • Static batching: the caller accumulates N items before sending; simple, but latency is uncontrolled;
  • Dynamic batching: the server merges requests arriving within a few-millisecond window into one batch and runs it on the GPU in a single pass — built into Triton, and it can boost small-model throughput 5-10x;
  • Continuous batching: LLM-specific, scheduling at token granularity so the GPU never sits idle; see LLM Inference Optimization.
text
Dynamic batching (non-LLM):
t=0ms  request A │
t=2ms  request B │──► batch(A,B,C) in one forward pass ──► three responses
t=4ms  request C │       (batch_timeout=10ms)

2. Result caching and merging similar requests ​

  • Exact caching: requests with an identical (model, input) return the cached result directly. In recommendation scenarios hit rates can reach 30-50% (hot items, repeated queries);
  • Similar-request merging (dedup): treat requests whose embedding distance is below a threshold as similar and return the cached representative — big gains, but handle with care (it affects correctness; set a business-level tolerance).

3. Concurrency model: async is king ​

  • Use async IO (asyncio) instead of "one thread per request + blocking calls";
  • Route inference calls through server-side batching wherever possible; don't let a single request monopolize a model invocation;
  • When multiple models share an instance, isolate them with per-model queues so one slow model can't drag everything down (see Multi-Model Serving with NVIDIA Triton).

4. The system layer: the easily overlooked 10-30% ​

  • Connection pool reuse (HTTP keep-alive / gRPC channel reuse) — avoid setting up a connection per request;
  • Serialization optimization: JSON → Protobuf, or message compression (gzip); significant gains on large payloads;
  • Async logging and metrics: never let synchronous logging block the request path;
  • NIC/kernel tuning: TCP BBR, buffer sizing. For cross-region deployments the network is often the biggest chunk of your P99.

4. Capacity Planning: From Estimation to Autoscaling ​

1. The core formula (Little's Law) ​

text
Concurrency (requests in flight) = QPS × average latency (seconds)

Example: target 1000 QPS at 50ms average latency → required concurrency = 1000 × 0.05 = 50

How much concurrency a single instance can handle → back out the instance count:

text
Instances = target concurrency / measured per-instance concurrency × 1.3 (30% safety margin)

2. The complete capacity planning workflow ​

text
(1) Load test: stress a single instance to plot its latency-throughput curve
    (the max QPS at which P99 still meets the target)
    (methodology: /practice/load-testing)
(2) Convert: target QPS / per-instance max QPS × 1.3 = instances needed
(3) Instrument: track live QPS, P99, and concurrency to validate the formula
(4) Autoscale: HPA on CPU or custom metrics; for GPU services prefer QPS or queue depth
(5) Reserve for peaks: size for peak QPS on big sale days (Double 11), not the average

Three rules for load testing

  1. Use a real traffic distribution (replay production traffic), not uniform synthetic traffic;
  2. Test against P99, not the average — P99 meeting target is what the SLA measures;
  3. Push until the point where "P99 starts to degrade" and record the QPS there — that's your capacity ceiling. Full methodology in Load Testing and Capacity Planning.

3. Elasticity strategies ​

Scaling dimensionMetricRecommendation
Horizontal scalingRequest queue depth / QPS / custom metricsBounded queues; scale up gradually, scale down quickly (avoid flapping)
GPU-specificvGPU or full cardsFor small models, partition cards with MIG/vGPU
Time-based planningCron-driven scaling preset from traffic curvesScale out 1 hour ahead of big sale events
text
Step 1: Load test to locate the bottleneck layer (don't guess)
Step 2: Model layer (quantization) — biggest gains, smallest change
Step 3: Service layer (batching, caching, async)
Step 4: Engine layer (TensorRT compilation) — deep specialization
Step 5: System layer (connection pools, network, IO)
Step 6: Capacity planning + elastic scaling — let resources follow traffic

After each step, load test again and quantify the gain; change one layer at a time, otherwise you won't be able to attribute the improvements.

Trade-offs ​

DecisionOptionsHow to choose
Latency vs throughputSmall batches protect latency / large batches boost throughputOnline: protect P99 latency; offline: max out throughput
Optimize vs add machinesTuning saves money but takes time / machines save time but cost moneyQuantize and tune the service layer first; scale only after that
P50 vs P99Optimize the median / optimize the tailAnchor the SLA on P99; chase the tail with dedicated investigations
Cache gains vs correctnessCaching is fast / freshness mattersDepends on the business's tolerance for stale results
Scale fast vs scale safeAggressive scale-in saves money / conservative avoids oscillationScale in slowly (prevent oscillation), scale out fast (prevent timeouts)

One-line summary: performance is measured, not tweaked — set metrics first, load test to locate the bottleneck, optimize layer by layer, size capacity with Little's Law, and make every dollar land on a measurable gain.

Further Reading ​

References ​