Skip to content

Deployment Design Principles

At a glance Inference deployment projects rarely fail because the engine isn't advanced enough — they die from design mistakes. This article distills twelve deployment design principles: latency-first vs throughput-first, memory budget, streaming beats waiting, three-layer bottleneck checks, activation outliers before quantization, speculative decoding vs batch size, 5% gray releases, KV hit-rate monitoring, failure drills, end-to-end benchmarks, and more.

Deployment Design Principles ​

The success or failure of an inference deployment project is decided before the first line of serving code is written.

Beginners tend to fixate on "which engine to use" — vLLM or TensorRT-LLM? AWQ or GPTQ? Speculative decoding or not? But real deployment postmortems tell us again and again: the engine is rarely the bottleneck; the design decisions are. Latency-first or throughput-first never decided → the entire tuning direction is wrong. Memory budget never calculated → OOM on day one. Streaming vs non-streaming never chosen → TTFT can never come down. No gray-release strategy → the first launch is an incident. NVIDIA keeps stressing in Deploying LLMs in Production:

"The stability of an inference service is not chosen with the engine — it is designed."

This article presents 12 actionable inference deployment design principles. Each answers three questions: why the principle matters, how to actually follow it, and what the most common counterexample looks like. They don't require you to master every engine first — on the contrary, most of these principles are written before the engine is even chosen. Read Deploy an Inference Service from Scratch to establish the overall flow, then let this article add "discipline" to every step.

Where this article sits

This is a "process layer" article: it is about how to organize an inference deployment project, not how to use a given engine. The list of what goes wrong lives in Common Pitfalls and Anti-Patterns; the specific parameters live in Tuning and Performance Optimization.

1. The Twelve Principles at a Glance ​

#PrincipleOne-linerTell-tale counter-sign
1Latency-first vs throughput-first: pick one firstYou can't have both; the binary choice is the starting point of all tuningWanting both p99 200 ms and 1000 tok/s on one GPU
2Calculate the memory budget firstWeights + KV cache + temporaries = total budgetOOM on day one
3Pre/post-processing in the same service as the modelFewer network hops, lower p99Tokenizer as a separate service → TTFT doubles
4Streaming beats waiting for the resultImproves first-token latency; experience engineeringChat API non-streaming → TTFT maxed out
5Check bottlenecks at all three layers: model, operator, systemDon't stare at GPU utilization aloneGPU at 100% but slow, because memory-bound
6Look at activation outlier distribution before quantizingModels with many outliers collapse under quantizationLLaMA-family accuracy drop with no idea why
7Speculative decoding depends on batch sizeLarge batch is a net lossEAGLE slower at large batch
8Gray-release at 5% traffic firstValidate new configs on small trafficFull rollout of AWQ, business breaks
9Monitoring is more than QPSAlso KV cache hit rate, temperatureQPS looks fine, p99 silently spiking
10Failure drills as routineDon't discover the missing runbook when it actually breakskill -9 one instance, whole site down
11Quantization is a trade-off, not a free lunchAccuracy loss is always there — measure itINT4 shipped, business quality silently degrades
12End-to-end benchmarks, not single-operator benchmarksUsers feel the end-to-endBeautiful single-operator benchmark, slow end-to-end

These 12 form a deployment loop:

        ┌───────────────────────────────────────────┐
        │   Business SLA sets direction (P1, P4)    │
        └───────────────────┬───────────────────────┘
                            ▼
        ┌───────────────────────────────────────────┐
        │   Memory budget and config (P2)           │
        └───────────────────┬───────────────────────┘
                            ▼
   ┌─────────────┬──────────┴───────────┬───────────────┐
   ▼             ▼                      ▼               ▼
 Engine tuning  Streaming design       Quantization    Speculative decoding
 (P5, P9)       (P3, P4)               decision (P6,  decision (P7)
                                        P11)
   └─────────────┴──────────┬───────────┴───────────────┘
                            ▼
        ┌───────────────────────────────────────────┐
        │   Gray release + monitoring (P8, P9, P10) │
        └───────────────────┬───────────────────────┘
                            ▼
        ┌───────────────────────────────────────────┐
        │   End-to-end benchmark verification (P12) │
        │   → back to tuning                         │
        └───────────────────────────────────────────┘

A common misunderstanding

"Design principles" sounds like theory preaching, but it is actually a time-saving methodology. Skip the memory budget and you may rework after OOM; skip streaming design and you may change the API only after users complain it's "frozen"; skip gray releases and you may write the runbook only after the incident postmortem. Principles are not constraints — they are the anchor that pulls you out of "incident on launch day."

2. Principle by Principle ​

Principle 1: Latency-first vs throughput-first — pick one first ​

Why. These two goals directly conflict in engineering: latency-first wants a small batch to protect p99; throughput-first wants a large batch to pull concurrency. Chasing both, you will keep raising max_num_seqs one moment and lowering it the next, never finding a stable configuration. The first job is to ask the business clearly: is a user waiting for the result, or is this a background batch job? That answer determines 90% of your tuning direction.

How. Pick one with this table:

Business formPrioritymax_num_seqsenable_chunked_prefillSpeculative decoding
Chat / code completionLatency-first16–32OnOn (pays off at small batch)
Document summarization / translationThroughput-first64–128OffOff (regresses at large batch)
Offline batchThroughput-first256+OffOff
RAG Q&ALatency-first (user waiting)16–32OnOn

Once defined, write it into the deployment doc, and check every tuning decision against "does my current direction still match the original definition?" See Tuning and Performance Optimization and Latency, Throughput, and Concurrency.

Counterexample. The classic one: to "save a bit of cost," a chat service raises max_num_seqs to 64 — p99 jumps from 200 ms to 1200 ms, users feel "it's frozen," complaints explode, and the team rolls back to 32. The rollback costs another week of tuning in other directions. The first decision was wrong, so every later tuning effort was pushing in the wrong direction. Second counterexample: not knowing what you want, you aim for "balance" — and both sides miss: latency misses the SLA and throughput misses capacity. Balance is not good; balance is failing both sides.

Principle 2: Calculate the memory budget first ​

Why. LLM inference memory has three layers: weights + KV cache + temporaries. Tuning without a calculated budget is blind tuning — you think there's room for 64 concurrency, and the KV pool blows up on the first run. This is the engineering application of The GPU Memory Hierarchy and the Bandwidth Wall.

How. Calculate the budget with the formula before deploying:

Total memory = weights + KV cache + temporaries

Weights: bf16 model = 2 × num_params bytes
         AWQ INT4 = 0.5 × num_params bytes

KV cache = 2 × num_layers × hidden_size × max_model_len × max_num_seqs × dtype_bytes

Temporaries: ≈ 1–2 GB (attention intermediates, kernel temporaries)

Example: Llama-2-7B bf16, max_model_len=4096, max_num_seqs=32
  Weights = 14 GB
  KV cache = 2 × 32 × 128 × 4096 × 32 × 2 bytes = 8 GB
  Temporaries = 1.5 GB
  Total = 23.5 GB  ← fits on an A100 40GB

Once the budget is calculated, every parameter change must answer "is it still within budget?" See the KV cache formula in Tuning and Performance Optimization.

Counterexample. Counterexample one: a team "goes by feel" with max_num_seqs=128 and OOMs on launch — KV cache takes 32 GB, weights 14 GB, temporaries 1.5 GB, total 47 GB over the A100's 40 GB. Counterexample two: the budget counts weights only, not KV cache — deploying a 70B model to a single A100 80GB, 140 GB of weights blows up outright — TP=2 is mandatory. Counterexample three: the budget is calculated to "exactly enough" with no temporary headroom — a long prompt's attention intermediates push it over the edge. Budget by "weights + KV + temporaries + 20% headroom."

Principle 3: Pre/post-processing in the same service as the model ​

Why. An inference service is not just the LLM — there is a tokenizer and prompt template in front, and a detokenizer, content filter, and formatting behind. Splitting these into standalone services adds two network hops per request, easily doubling p99. The cost of network hops is especially large in LLM scenarios — inference itself takes hundreds of ms, so an extra 50 ms of network doesn't look like much, but it more than doubles the p99 (see queuing theory).

How. Put the tokenizer, prompt template, detokenizer, and content filter in the same service process:

✓ Recommended architecture:

  HTTP request ─→ FastAPI ─→ tokenizer ─→ vLLM ─→ detokenizer ─→ content filter ─→ HTTP response
                 (one process)                                  (one process)


✗ Anti-pattern:

  HTTP request ─→ tokenizer svc ─→ HTTP ─→ vLLM ─→ HTTP ─→ detokenizer svc ─→ HTTP ─→ content filter svc ─→ HTTP
  (5 network hops, p99 doubles)

The only exception is independent models like embedders / rerankers — they must be separate services (different weights), but still keep them on the same host to reduce network overhead. See Model Serving and Orchestration.

Counterexample. Counterexample one: microservice dogma splits the tokenizer into its own service "for reusability" — p99 doubles, and the reusability never pays off. Counterexample two: the content filter runs on the client, but client implementations drift, and production filtering behavior diverges. Counterexample three: post-processing (e.g., JSON parsing) is pushed to an async queue "to not block the response" — the user sees a response without the structured fields and must poll for the result, which is another full round trip.

Principle 4: Streaming beats waiting for the result ​

Why. The user's sense of "frozen" is driven by time to first token (TTFT), not end-to-end latency. For the same 5-second generation, streaming shows the first character at 200 ms and feels like "it's thinking," while non-streaming shows the complete result after 5 seconds and feels like "it's dead." This is the core of experience engineering — perceived time ≠ actual generation time.

How. Online services (chat, completion) default to streaming, using SSE (Server-Sent Events):

GET /v1/chat/completions
Content-Type: text/event-stream

data: {"choices":[{"delta":{"content":"Paged"}}]}
data: {"choices":[{"delta":{"content":"Attention"}}]}
data: {"choices":[{"delta":{"content":"is"}}]}
data: {"choices":[{"delta":{"content":"vLLM"}}]}
data: [DONE]

No WebSocket (bidirectional, complex), no polling (a request per second is N extra requests). See the "streaming chat API" project in Portfolio Projects.

Counterexample. Counterexample one: the chat API is designed as "wait for the model to finish, return everything at once" — TTFT is 5 seconds and the user is gone. Counterexample two: polling simulates streaming (the frontend requests increments every 200 ms) — every poll is a full HTTP round trip, QPS multiplies by orders of magnitude, and the service gets hammered. Counterexample three: WebSocket for streaming — the bidirectional channel is never used, and complexity goes up. SSE is the standard answer for streaming output.

Principle 5: Check bottlenecks at all three layers — model, operator, system ​

Why. An inference service's bottleneck can sit at three layers: the model layer (too many FLOPs), the operator layer (poor kernel implementation), and the system layer (scheduling, batching, network). GPU utilization alone cannot tell you which — 100% GPU can be compute-bound (model layer) or memory-bound (operator layer). See The Roofline Model and Compute Analysis and GPU Architecture and Optimization.

How. Use the four-step diagnosis from Tuning and Performance Optimization:

Step 1: nvidia-smi — GPU utilization, memory utilization, temperature
Step 2: vLLM logs — KV cache usage, throughput, swapped
Step 3: Nsight Systems — operator time distribution
Step 4: Locate the bottleneck layer:
       - Model layer → quantize, distill, prune
       - Operator layer → FlashAttention, kernel fusion
       - System layer → tune max_num_seqs, add PagedAttention

The remedies are completely different per layer — blind tuning is straining at the wrong layer.

Counterexample. Counterexample one: GPU utilization is 100% but slow; the team assumes insufficient compute and adds GPUs — it was memory-bound, and adding GPUs doesn't fix memory bandwidth. Counterexample two: single-operator benchmarks look great (attention 5 ms) but end-to-end is slow — the bottleneck is the scheduling layer (kernel launches too frequent), and more operator optimization won't help. Counterexample three: only throughput is watched, never latency — p99 silently spikes, unnoticed. See Principle 9.

Principle 6: Look at activation outlier distribution before quantizing ​

Why. Quantization compresses continuous values into discrete buckets — models with a normal (bell-shaped) activation distribution quantize losslessly, while models with many activation outliers (like the "massive activations" of the LLaMA family) collapse. Quantizing without checking the activation distribution is gambling. This is the core trade-off of Model Quantization Fundamentals and the root motivation of Weight-Only Quantization and Mixed Precision — why is weight-only quantization more robust than weight+activation quantization in LLM scenarios? Because activation outliers are far more severe than weight outliers.

How. Run an activation distribution check before quantizing:

python
# Run a few dozen samples through transformers; statistics per layer
from transformers import AutoModel
model = AutoModel.from_pretrained("meta-llama/Llama-2-7b-chat-hf")

# Hook each layer's activations while running samples
activations = {}
def hook(name):
    def fn(module, input, output):
        activations[name] = output.detach()
    return fn

for name, module in model.named_modules():
    if "mlp.gate" in name or "attention" in name:
        module.register_forward_hook(hook(name))

# After running samples, check each layer's max/mean ratio
for name, act in activations.items():
    ratio = act.max() / act.mean()
    print(f"{name}: max/mean = {ratio:.1f}")  # > 100 means many outliers

For layers with ratio > 100, be careful with quantization — consider weight-only (keep activations at full precision), or activation-aware schemes like AWQ/GPTQ. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

Counterexample. Counterexample one: a team runs SmoothQuant (activation quantization) on LLaMA-2 and MMLU drops 3 points — LLaMA's activation outliers got crushed. Counterexample two: INT8 weight + INT8 activation on LLaMA runs slower than bf16 — dequantization overhead exceeds the memory savings (see "INT8 weights + FP16 activations compute slower" in Common Pitfalls and Anti-Patterns). Counterexample three: AWQ ships without a business evaluation — business quality silently degrades: MMLU barely moves, but the business prompt distribution drops 5 points.

Principle 7: Speculative decoding depends on batch size ​

Why. Speculative decoding gains depend heavily on batch size — at small batch, the large model's compute sits idle and every correct small-model guess is free; at large batch, the large model is already saturated, and the draft model only adds memory and compute. This is the core trade-off of Speculative Decoding and Medusa/EAGLE, and one of the most common anti-patterns.

How. The "golden range" of speculative decoding is batch ≤ 16 — beyond 32 is the break-even point, beyond 64 a net loss. In production:

ScenarioBatch sizeSpeculative decodingRationale
Online chat1–16OnGains are largest at small batch
Online batch32–64OffRegresses at large batch
Offline batch256+OffPure net loss

If your business has both "small-batch online" and "large-batch offline" traffic, split the traffic with a router — speculative-decoding instances take the small-batch traffic, regular instances take the large-batch traffic. See the "multi-model routing" project in Portfolio Projects.

Counterexample. Counterexample one: EAGLE enabled for a large-batch offline scenario — throughput regresses 30% — the large model is already compute-saturated, and the draft model just occupies memory. Counterexample two: online chat without speculative decoding — TTFT can't go below 50 ms — small-batch is the sweet spot; not enabling it wastes free performance. Counterexample three: unaware of the batch dependency, you set max_num_seqs to 64 with EAGLE on, then debug for a week without finding the cause — it is a design conflict at the root.

Principle 8: Gray-release at 5% traffic first ​

Why. New configurations (new quantization, new engine, new parameters) always have unexpected side effects — accuracy drift, long-tail latency, memory leaks, client compatibility. Full rollout = exposing everyone to the problem. A gray release contains the problem within 5% of traffic: small blast radius, fast rollback. This is standard practice in Model Serving and Orchestration and the canonical use of the multi-version capability in Triton Inference Server.

How. The standard gray-release flow:

Day 0: Deploy the new config to 1 instance (5% of instances)
       - Traffic split 5:95 new:old by weight
       - Watch QPS / p99 / error rate / business metrics for 24 hours
Day 1: No anomalies → expand to 20%
       - Another 24 hours of observation
Day 2: No anomalies → expand to 50%
       - From here a problem hurts more but stays controllable
Day 3: No anomalies → 100% full rollout
       - Decommission the old config

Every step needs an automatic rollback mechanism: p99 beyond 1.5× SLA → auto-rollback; error rate > 1% → auto-rollback; business metrics drop 5% → auto-rollback. See "gray release and rollback" in Model Serving and Orchestration.

Counterexample. Counterexample one: full rollout of AWQ INT4, business quality collapses — MMLU unchanged, but the business prompt distribution drops 10 points. Rollback takes an hour; the business suffers in the meantime. Counterexample two: the new version is tested for latency only, not business metrics — latency looks fine, but the output format changed (INT4 shifted the logits), and client parsing fails. Counterexample three: no rollback plan prepared — when the problem surfaces, the team scrambles; the rollback takes longer than the problem itself.

Principle 9: Monitoring is more than QPS ​

Why. QPS is "traffic," not "health" — QPS can be normal while p99 silently spikes, KV cache hit rate drops, or GPU thermal throttling kicks in. These are all "looks fine, actually collapsing" signals. Production monitoring must be multi-dimensional, each with its own alert threshold.

How. The minimal monitoring set for an inference service:

MetricMeaningAlert threshold
QPSTrafficSudden change > 50%
p50 / p99 latencyPerceived qualityp99 > 1.5× SLA
GPU utilizationCompute usage< 50% sustained for 5 minutes
GPU memory utilizationMemory usage> 95%
GPU temperatureCooling> 85°C
KV cache hit ratePrefix reuse< 50%
Swapped request countMemory swap-out> 0
Error rateAnomalies> 1%
Business metricsEnd-to-endDrop of 5%

The three bold items are unique to LLM inference and most often ignored. See "vLLM logs under the microscope" in Tuning and Performance Optimization.

Counterexample. Counterexample one: only QPS and error rate monitored — p99 silently climbs from 200 ms to 800 ms, discovered only via user complaints. Counterexample two: GPU utilization monitored without temperature — thermal throttling silently costs 30% of performance, unnoticed until users ask "why is it slower?" Counterexample three: KV cache hit rate ignored — in a multi-turn scenario, prefix caching silently fails and TTFT doubles.

Principle 10: Failure drills as routine ​

Why. Failure is not "if" but "when." Production will meet: instance crashes, GPU failures, network partitions, full disks, dependency timeouts. Without drills, the failure finds you unprepared — you say "we have redundancy," and then discover failover was never configured; you say "we have monitoring," and then discover the alert thresholds were wrong. Netflix's Chaos Monkey thinking applies equally to inference services: break things on purpose and verify the system's resilience.

How. Drill at least four failure types:

Failure typeDrill methodWhat to verify
Single-instance crashkubectl delete pod — kill a random podDoes traffic shift automatically? Does p99 stay within SLA?
GPU failurenvidia-smi -r — reset the GPUDoes the instance restart automatically? Does the alert fire?
Network partitioniptables simulating packet lossDoes the client retry? Does a retry storm (cascading failure) form?
Upstream timeoutSimulate 5-second LLM responsesIs degradation triggered (route to the small model or return cache)?

Drill monthly; write a postmortem each time. See "failure drills" in Model Serving and Orchestration.

Counterexample. Counterexample one: no drills — when a GPU fails, the instance restarts automatically but the KV cache is lost; every in-flight request fails and client retry storms cascade. Counterexample two: during a drill, you discover failover was misconfigured — killing a pod returns 503s to users instead of shifting traffic. Counterexample three: no degradation for upstream timeouts — a single LLM timeout stalls the entire request chain.

Principle 11: Quantization is a trade-off, not a free lunch ​

Why. Quantization is one of the most effective inference optimization techniques, but it always has a cost — accuracy loss, tuning complexity, engine compatibility. Every team that treated quantization as a silver bullet has been burned: INT4 shipped, business quality silently degrades; MMLU unchanged but JSON output formats break; quantization results inconsistent across engines. See "the trade-offs of quantization" in Model Quantization Fundamentals.

How. A quantization decision must answer four questions:

Q1: How tolerant is the business of accuracy loss?
   High (chat, summarization) → AWQ INT4 is an option
   Low (math, code)          → stay bf16, INT8 at most

Q2: What is the model's activation outlier distribution?
   Normal        → weight+activation quantization is viable
   Many outliers → weight-only quantization only

Q3: Has the business evaluation set been run before and after?
   Run, < 1 point drop → ship
   Not run             → do not ship

Q4: Are the quantization costs accounted?
   Gains: 1.5–2× throughput, 1/4 memory
   Costs: accuracy loss, engine compatibility, debugging complexity

See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

Counterexample. Counterexample one: INT4 shipped without a business evaluation — MMLU barely moves but the business prompt distribution drops 5 points. Counterexample two: activation quantization without checking activation distribution — LLaMA outliers crushed, output format changes. Counterexample three: assuming quantization always speeds things up — but with a small model + large batch, dequantization overhead exceeds the memory savings, and performance actually regresses.

Principle 12: End-to-end benchmarks, not single-operator benchmarks ​

Why. A single-operator benchmark (e.g., "attention in 5 ms") reflects engine implementation quality but cannot reflect user experience. What the user feels is end-to-end: HTTP request → tokenizer → prefill → decode → detokenizer → content filter → HTTP response. Attention is just one segment; the bottleneck can be entirely elsewhere. Making decisions with single-operator benchmarks is viewing a panorama through a microscope.

How. Deployment decisions must be based on end-to-end benchmarks:

python
# End-to-end latency (full HTTP request → complete response)
import time, requests
t0 = time.perf_counter()
resp = requests.post("http://llm-service/v1/chat/completions",
                    json={"messages": [...]})
t1 = time.perf_counter()
print(f"End-to-end: {(t1-t0)*1000:.0f} ms")

Report both "end-to-end latency" and "in-engine latency" (measured by vLLM itself) — the former is user perception, the latter engine capability. See "what to measure" in Inference Benchmarking in Practice.

Counterexample. Counterexample one: the team optimizes the attention operator from 5 ms to 3 ms, but end-to-end latency doesn't move — the bottleneck is HTTP and tokenization. Counterexample two: single-operator benchmarks look great, but after launch, poor scheduling makes end-to-end twice as slow. Counterexample three: when switching engines, only in-engine benchmarks are compared, never end-to-end — the new engine is faster internally but slower in the HTTP layer, a net regression.

3. Tensions Between Principles ​

The 12 principles are not independent — they pull against each other and must be balanced per scenario:

Tension 1: Principle 1 (latency vs throughput) vs Principle 8 (gray release) ​

A latency-first config (small batch) may look fine during gray release because "traffic is small," surfacing problems only at full rollout. Response: during gray release, deliberately overdrive traffic (5% of instances carrying 20% of traffic) to stress-test.

Tension 2: Principle 4 (streaming) vs Principle 9 (monitoring) ​

Streaming latency distributions differ completely from non-streaming — you cannot reuse the same monitoring thresholds. Response: for streaming services, p99 must be chunk-level latency, not request-response latency. See "streaming test methods" in Inference Benchmarking in Practice.

Tension 3: Principle 6 (activation outliers) vs Principle 11 (quantization is a trade-off) ​

A model with many activation outliers goes weight-only, but weight-only accelerates less than weight+activation on some engines. Response: check business accuracy tolerance first, then pick the scheme — in accuracy-sensitive scenarios, keep weight-only even if slower.

Tension 4: Principle 7 (speculative decoding vs batch) vs Principle 1 (latency vs throughput) ​

Latency-first (small batch) happens to be the sweet spot of speculative decoding, while throughput-first (large batch) means turning it off. Response: the two don't actually conflict — latency-first picks small batch + speculative decoding, throughput-first picks large batch + no speculative decoding. They are two sides of the same decision.

4. Turning Principles into Team Norms ​

The 12 principles only bind when codified into team norms. Suggested structure:

markdown
# Team Deployment Standards

## 1. Pre-deployment checklist
- [ ] Business SLA defined (P1)
- [ ] Memory budget calculated (P2)
- [ ] Streaming / non-streaming decided (P4)
- [ ] Quantization scheme evaluated (P6, P11)
- [ ] Speculative decoding applicability evaluated (P7)

## 2. During-deployment norms
- [ ] Pre/post-processing in the same service as the model (P3)
- [ ] Three-layer bottleneck diagnosis completed (P5)
- [ ] End-to-end benchmark measured (P12)

## 3. Launch norms
- [ ] Gray-release process (P8)
- [ ] Monitoring and alert thresholds (P9)
- [ ] Failure drill plan (P10)

## 4. Operations norms
- [ ] Monthly failure drills
- [ ] Quarterly SLA review
- [ ] Postmortem for every incident

See "team standards template" in Model Serving and Orchestration.

5. Further Reading ​

References ​