Skip to content

Load Testing and Capacity Planning

At a glance Three questions you must answer before launch: how much QPS can the service take, what is the P99, and does it need scaling. This guide covers a full load testing methodology: tool selection (wrk/locust/vegeta/ghz/k6), scenario design, interpreting results, and capacity formulas.

Load testing is the process of using controlled traffic to expose a service's real capability limits; capacity planning is the math problem of using load test data to answer "how many machines do we need." Together they answer the three questions that must be settled before launch: How much QPS can we handle? What is the P99 latency? How many instances do we need? Going live without load test data is running an unprotected experiment on real user traffic — capacity estimates off by 10x, P99 quietly degrading, a promo day knocking the service flat: these are the classic incidents of untested deployments. This guide gives you a complete methodology, from goal setting and tool selection through capacity formulas, so that when your hand reaches for the "launch" button you have an evidence-backed answer sheet. For the theory, see Performance Metrics and Tuning and Monitoring.

Three Numbers to Internalize First

Three basics of load testing: measure P99, not averages (averages hide the tail); always warm up (data from a cold service is skewed); always record the environment (a load test you cannot reproduce never happened).

1. Load Testing Goals: Answer Three Questions at Once ​

A proper load test produces three conclusions:

QuestionHow it is answeredDeliverable
How much QPS can it take?Step up the load until you find the throughput kneeSaturation QPS, max concurrency
What is the P99?Latency distribution at each concurrency levelLatency-vs-concurrency curve
Do we need to scale?Plug peak traffic into the capacity formulaInstance count, GPU count

Missing any of the three, the report is half-finished. "How much QPS" and "what is the P99" must be answered together — isolated numbers lie: the service may already be at 500ms P99 when it hits 500 QPS (latency degrades first), or it may still be rock solid at 2000 QPS (capacity to spare). It is the curve formed by these two dimensions that justifies capacity decisions.

2. Tool Selection: Five Cards and When to Play Them ​

ToolPositioningBest ForLearning CurveKey Features
wrkSingle-machine HTTP benchmarkQuickly gauging raw throughput, hammering /healthzLow (command line)Multithreaded + connection reuse; weak scripting
vegetaSingle-machine HTTP at a target QPSFixed-RPS attacks, targets stored as filesLowPrecise attack-rate control, good output formats
locustDistributed scenario scriptingBusiness paths (login → query → checkout), with think timeMedium (Python)Programmable, master-worker distribution, web UI
ghzgRPC load testinggRPC servicesMediumNative proto support, streaming calls
k6Cloud-native/CI ecosystemK8s environments, pipeline integration, SLO assertionsMedium (JS scripts)Built-in assertions and thresholds, K8s Operator

The one-line decision rule: just probing raw throughput → wrk; simulating realistic business rhythms → locust; gRPC service → ghz; SLO assertions in CI → k6; hitting an exact fixed QPS → vegeta. For more tools, see Tools and Resources.

bash
# wrk: probe raw throughput (4 threads, 100 connections, 10 seconds)
wrk -t4 -c100 -d10s http://localhost:8000/healthz

# vegeta: hit exactly 200 QPS for 30 seconds, output a histogram
echo "POST http://localhost:8000/predict" | vegeta attack -rate=200 -duration=30s \
  -body=payload.json | vegeta report --type=hist[0,50ms,100ms,200ms,500ms]

3. The Load Testing Workflow: Five Steps to a Capacity Decision ​

Step 1: Define target metrics (write the SLO before you test) ​

Write down "what counts as passing" before you run anything, or you won't be able to draw conclusions afterwards:

text
Draft SLO (example):
- Throughput: steadily sustain 200 QPS (2x the current estimated production peak)
- Latency: P99 < 100ms, P50 < 30ms
- Stability: error rate < 0.1%, no 5xx spikes
- Resources: per-instance CPU < 80%, GPU utilization < 85% (15% headroom)

For full SLO design (burn rates, alert thresholds), see Putting Observability into Practice and SLO concepts.

Step 2: Scenario design — four load patterns that cover reality ​

ScenarioHowQuestion it answers
Fixed QPSvegeta/k6 at a fixed rate for 10 minutesSteady state: is it stable under this load?
Step loadlocust doubling concurrency step by stepWhere is the throughput knee? Where does P99 start degrading?
Peak spikeSudden jump from 100 to 300 QPS held for 60sCan it survive bursts? Does it recover quickly?
Mixed workload90% normal traffic + 10% slow requests/heavy modelDo slow requests clog the queue (head-of-line blocking)?

Step loading is the key to finding the knee: start at 50 concurrent users and double every 60 seconds (50 → 100 → 200 → 400), recording QPS and P99 along the way. The point where QPS stops growing with concurrency is the throughput knee (saturation point).

python
# locustfile.py — step load + mixed workload skeleton
from locust import HttpUser, task, between
import random

class MixedUser(HttpUser):
    wait_time = between(0.05, 0.2)     # simulate realistic request gaps
    @task(9)
    def predict_small(self):           # 90% light requests
        self.client.post("/predict", files=self._img("small"))
    @task(1)
    def predict_large(self):           # 10% heavy requests
        self.client.post("/predict", files=self._img("large"))

Step 3: Execution and monitoring — watch client and server together ​

While the test runs, collect data from both ends; looking only at the client mislocates the bottleneck:

  • Client (load tool output): QPS, latency distribution, error rate;
  • Server (Prometheus): CPU, memory, GPU utilization, queue depth, GC/memory watermark;
  • Cross-check: client QPS climbs but server CPU sits at 30% → the bottleneck is the network or the load generator, not the service.

For GPU workloads there are two more numbers to watch: GPU utilization and GPU memory usage (nvidia-smi or DCGM; collection methods in Putting Observability into Practice). GPU pegged while CPU idles means batching isn't filling up; GPU idle while CPU is pegged means preprocessing or deserialization is the bottleneck.

Step 4: Reading the numbers — four signals to understand ​

text
    QPS
    │        ┌── knee: throughput caps out, more load only adds latency
    │       /
    │      /
    │     /
    │────/───────────────────── latency degradation zone
    │   /
    │  /
    │ /
    └──────────────────────────▶ concurrency

Four key signals:

  1. Throughput knee: QPS stops growing with concurrency. Past the knee you are "trading latency for QPS" — stop.
  2. P99 drifts with concurrency: P99 stable at low concurrency, then degrading roughly linearly as concurrency rises — provision production capacity before the P99 degradation point.
  3. Error rate lifts off: a jump from 0 to 0.1% to 1% is usually the precursor of timeouts, connection-pool exhaustion, or OOM.
  4. GPU-bound vs CPU-bound: GPU pegged → switch engine, quantize, or increase batch size; CPU pegged → optimize preprocessing, add workers, move to GPU or bigger machines.

Step 5: Capacity math — two formulas and one coefficient ​

text
concurrency ≈ QPS × P99 latency (seconds)          # practical form of Little's Law
peak QPS ≈ daily average QPS × peak coefficient (typically 3-10x)
instances needed = peak QPS ÷ safe per-instance QPS (70-80% of the knee)

A worked example:

text
20M requests/day → daily average QPS ≈ 231 → peak coefficient of 5 → peak QPS ≈ 1155
Load test finds a per-instance knee of 400 QPS → safe capacity at 70% = 280 QPS
Instances needed = ceil(1155 / 280) = 5

Three corrections people forget: the peak coefficient must match your actual business (e-commerce mega-sales go well beyond 10x; internal SaaS tools run 2-3x); reserve 20-30% headroom for forecast error; multiply by 2 for multi-region/DR. Hardware differences directly change capacity — see Hardware for selection dimensions.

Don't Apply the Concurrency Formula Backwards

A common mistake is conflating "concurrency" with "QPS". concurrency = QPS × latency: at 100ms latency, targeting 1000 QPS → about 100 in-flight requests suffice; if the same service's latency rises to 500ms, the same 1000 QPS needs 500 in flight — latency degradation automatically amplifies concurrency demand, which is one mechanism behind peak-hour avalanches.

4. Load Test Report Template ​

Results must be reproducible and trustworthy — to others, or to you two weeks from now. Use this template:

markdown
# Load Test Report: image-classifier v1

## Environment
- Load generator: 8C16G VM, wrk 4.6 / locust 2.30
- Service host: 4C8G, Docker, 4 workers, onnxruntime 1.19.2
- Model: mobilenetv2 ONNX INT8, input (1,3,32,32)

## Target SLO
- 200 QPS sustained, P99<100ms, error rate<0.1%

## Results
| Concurrency | QPS | P50(ms) | P99(ms) | Errors | CPU% | Notes |
| --- | --- | --- | --- | --- | --- | --- |
| 50 | 310 | 12 | 38 | 0% | 55% | steady state |
| 100 | 410 | 18 | 72 | 0% | 78% | near the knee |
| 200 | 430 | 45 | 380 | 0.3% | 92% | past the knee, P99 degrading |

## Conclusions
- Safe per-instance capacity ≈ 280 QPS (65% of the 430 knee)
- Estimated production peak 1155 QPS → need 5 instances (incl. 1 spare)
- Risks: P99 degrades fast past 100 concurrency; tune batching and re-test first

5. Common Pitfalls: Why Load Test Numbers Can't Be Trusted ​

  1. The load generator becomes the bottleneck: 8-thread wrk can't saturate a 4-worker service because the generator's CPU pegs first. Fix: check the generator's load with top; if you can't push harder, add threads or go distributed (locust master-worker).
  2. Insufficient warm-up: a freshly started service with cold JIT/caches produces numbers that are too low or too high for the first tens of seconds. Fix: warm up for 30-60 seconds before recording anything.
  3. Inflated numbers from caching: every request hits the cache, you measure 5000 QPS, and real traffic with zero cache hits collapses the service. Fix: use unique parameters or disable caching, and test the "true path" separately.
  4. No baseline: testing the new version without ever testing the old one leaves you unable to say whether things got faster or slower. Fix: run the same script in the same environment before and after optimization (see the baseline table in Model Optimization in Practice).
  5. Testing for only 30 seconds: memory leaks and connection leaks need 10+ minutes to surface. Fix: run steady-state scenarios for at least 10 minutes while watching the server's memory curve (see pitfall #4 in Memory Leaks and OOM).

6. After the Test: Turn Data into Decisions ​

Load testing isn't about producing reports; it exists to drive three decisions: launch (SLO met with margin to spare), scale (SLO missed but more machines can fix it), or optimize (SLO missed and machines won't help — for instance, latency itself is over budget). The decision path:

text
Load test result ──▶ SLO met? ──yes──▶ Launch (keep monitoring; see /practice/observability-practice)
               │
               └──no──▶ Capacity problem or latency problem?
                           ├─ Capacity (QPS shortfall) ──▶ Scale out / add workers / quantize
                           └─ Latency (P99 over budget) ──▶ Optimize preprocessing / switch engine / improve batching

After any capacity change, re-test to verify — formulas only estimate; measurements answer.

Checklist ​

  • [ ] SLO written before testing (QPS, P99, and error rate, all three);
  • [ ] Tool matches protocol and scenario (HTTP/gRPC/distributed/CI);
  • [ ] Warm-up ≥ 30 seconds, steady-state runs ≥ 10 minutes;
  • [ ] Client and server metrics collected together (including GPU utilization);
  • [ ] Throughput knee identified, with P99 recorded before and after it;
  • [ ] Instance count computed via the capacity formula, with a safety margin stated;
  • [ ] Report contains all four parts: environment, parameters, results table, conclusions;
  • [ ] Load generator CPU never became the bottleneck (record the generator's load).

Further Reading ​

References ​