Appearance
Deployment Design Principles
The success or failure of an inference deployment project is decided before the first line of serving code is written.
Beginners tend to fixate on "which engine to use" — vLLM or TensorRT-LLM? AWQ or GPTQ? Speculative decoding or not? But real deployment postmortems tell us again and again: the engine is rarely the bottleneck; the design decisions are. Latency-first or throughput-first never decided → the entire tuning direction is wrong. Memory budget never calculated → OOM on day one. Streaming vs non-streaming never chosen → TTFT can never come down. No gray-release strategy → the first launch is an incident. NVIDIA keeps stressing in Deploying LLMs in Production:
"The stability of an inference service is not chosen with the engine — it is designed."
This article presents 12 actionable inference deployment design principles. Each answers three questions: why the principle matters, how to actually follow it, and what the most common counterexample looks like. They don't require you to master every engine first — on the contrary, most of these principles are written before the engine is even chosen. Read Deploy an Inference Service from Scratch to establish the overall flow, then let this article add "discipline" to every step.
Where this article sits
This is a "process layer" article: it is about how to organize an inference deployment project, not how to use a given engine. The list of what goes wrong lives in Common Pitfalls and Anti-Patterns; the specific parameters live in Tuning and Performance Optimization.
1. The Twelve Principles at a Glance
| # | Principle | One-liner | Tell-tale counter-sign |
|---|---|---|---|
| 1 | Latency-first vs throughput-first: pick one first | You can't have both; the binary choice is the starting point of all tuning | Wanting both p99 200 ms and 1000 tok/s on one GPU |
| 2 | Calculate the memory budget first | Weights + KV cache + temporaries = total budget | OOM on day one |
| 3 | Pre/post-processing in the same service as the model | Fewer network hops, lower p99 | Tokenizer as a separate service → TTFT doubles |
| 4 | Streaming beats waiting for the result | Improves first-token latency; experience engineering | Chat API non-streaming → TTFT maxed out |
| 5 | Check bottlenecks at all three layers: model, operator, system | Don't stare at GPU utilization alone | GPU at 100% but slow, because memory-bound |
| 6 | Look at activation outlier distribution before quantizing | Models with many outliers collapse under quantization | LLaMA-family accuracy drop with no idea why |
| 7 | Speculative decoding depends on batch size | Large batch is a net loss | EAGLE slower at large batch |
| 8 | Gray-release at 5% traffic first | Validate new configs on small traffic | Full rollout of AWQ, business breaks |
| 9 | Monitoring is more than QPS | Also KV cache hit rate, temperature | QPS looks fine, p99 silently spiking |
| 10 | Failure drills as routine | Don't discover the missing runbook when it actually breaks | kill -9 one instance, whole site down |
| 11 | Quantization is a trade-off, not a free lunch | Accuracy loss is always there — measure it | INT4 shipped, business quality silently degrades |
| 12 | End-to-end benchmarks, not single-operator benchmarks | Users feel the end-to-end | Beautiful single-operator benchmark, slow end-to-end |
These 12 form a deployment loop:
┌───────────────────────────────────────────┐
│ Business SLA sets direction (P1, P4) │
└───────────────────┬───────────────────────┘
▼
┌───────────────────────────────────────────┐
│ Memory budget and config (P2) │
└───────────────────┬───────────────────────┘
▼
┌─────────────┬──────────┴───────────┬───────────────┐
▼ ▼ ▼ ▼
Engine tuning Streaming design Quantization Speculative decoding
(P5, P9) (P3, P4) decision (P6, decision (P7)
P11)
└─────────────┴──────────┬───────────┴───────────────┘
▼
┌───────────────────────────────────────────┐
│ Gray release + monitoring (P8, P9, P10) │
└───────────────────┬───────────────────────┘
▼
┌───────────────────────────────────────────┐
│ End-to-end benchmark verification (P12) │
│ → back to tuning │
└───────────────────────────────────────────┘A common misunderstanding
"Design principles" sounds like theory preaching, but it is actually a time-saving methodology. Skip the memory budget and you may rework after OOM; skip streaming design and you may change the API only after users complain it's "frozen"; skip gray releases and you may write the runbook only after the incident postmortem. Principles are not constraints — they are the anchor that pulls you out of "incident on launch day."
2. Principle by Principle
Principle 1: Latency-first vs throughput-first — pick one first
Why. These two goals directly conflict in engineering: latency-first wants a small batch to protect p99; throughput-first wants a large batch to pull concurrency. Chasing both, you will keep raising max_num_seqs one moment and lowering it the next, never finding a stable configuration. The first job is to ask the business clearly: is a user waiting for the result, or is this a background batch job? That answer determines 90% of your tuning direction.
How. Pick one with this table:
| Business form | Priority | max_num_seqs | enable_chunked_prefill | Speculative decoding |
|---|---|---|---|---|
| Chat / code completion | Latency-first | 16–32 | On | On (pays off at small batch) |
| Document summarization / translation | Throughput-first | 64–128 | Off | Off (regresses at large batch) |
| Offline batch | Throughput-first | 256+ | Off | Off |
| RAG Q&A | Latency-first (user waiting) | 16–32 | On | On |
Once defined, write it into the deployment doc, and check every tuning decision against "does my current direction still match the original definition?" See Tuning and Performance Optimization and Latency, Throughput, and Concurrency.
Counterexample. The classic one: to "save a bit of cost," a chat service raises max_num_seqs to 64 — p99 jumps from 200 ms to 1200 ms, users feel "it's frozen," complaints explode, and the team rolls back to 32. The rollback costs another week of tuning in other directions. The first decision was wrong, so every later tuning effort was pushing in the wrong direction. Second counterexample: not knowing what you want, you aim for "balance" — and both sides miss: latency misses the SLA and throughput misses capacity. Balance is not good; balance is failing both sides.
Principle 2: Calculate the memory budget first
Why. LLM inference memory has three layers: weights + KV cache + temporaries. Tuning without a calculated budget is blind tuning — you think there's room for 64 concurrency, and the KV pool blows up on the first run. This is the engineering application of The GPU Memory Hierarchy and the Bandwidth Wall.
How. Calculate the budget with the formula before deploying:
Total memory = weights + KV cache + temporaries
Weights: bf16 model = 2 × num_params bytes
AWQ INT4 = 0.5 × num_params bytes
KV cache = 2 × num_layers × hidden_size × max_model_len × max_num_seqs × dtype_bytes
Temporaries: ≈ 1–2 GB (attention intermediates, kernel temporaries)
Example: Llama-2-7B bf16, max_model_len=4096, max_num_seqs=32
Weights = 14 GB
KV cache = 2 × 32 × 128 × 4096 × 32 × 2 bytes = 8 GB
Temporaries = 1.5 GB
Total = 23.5 GB ← fits on an A100 40GBOnce the budget is calculated, every parameter change must answer "is it still within budget?" See the KV cache formula in Tuning and Performance Optimization.
Counterexample. Counterexample one: a team "goes by feel" with max_num_seqs=128 and OOMs on launch — KV cache takes 32 GB, weights 14 GB, temporaries 1.5 GB, total 47 GB over the A100's 40 GB. Counterexample two: the budget counts weights only, not KV cache — deploying a 70B model to a single A100 80GB, 140 GB of weights blows up outright — TP=2 is mandatory. Counterexample three: the budget is calculated to "exactly enough" with no temporary headroom — a long prompt's attention intermediates push it over the edge. Budget by "weights + KV + temporaries + 20% headroom."
Principle 3: Pre/post-processing in the same service as the model
Why. An inference service is not just the LLM — there is a tokenizer and prompt template in front, and a detokenizer, content filter, and formatting behind. Splitting these into standalone services adds two network hops per request, easily doubling p99. The cost of network hops is especially large in LLM scenarios — inference itself takes hundreds of ms, so an extra 50 ms of network doesn't look like much, but it more than doubles the p99 (see queuing theory).
How. Put the tokenizer, prompt template, detokenizer, and content filter in the same service process:
✓ Recommended architecture:
HTTP request ─→ FastAPI ─→ tokenizer ─→ vLLM ─→ detokenizer ─→ content filter ─→ HTTP response
(one process) (one process)
✗ Anti-pattern:
HTTP request ─→ tokenizer svc ─→ HTTP ─→ vLLM ─→ HTTP ─→ detokenizer svc ─→ HTTP ─→ content filter svc ─→ HTTP
(5 network hops, p99 doubles)The only exception is independent models like embedders / rerankers — they must be separate services (different weights), but still keep them on the same host to reduce network overhead. See Model Serving and Orchestration.
Counterexample. Counterexample one: microservice dogma splits the tokenizer into its own service "for reusability" — p99 doubles, and the reusability never pays off. Counterexample two: the content filter runs on the client, but client implementations drift, and production filtering behavior diverges. Counterexample three: post-processing (e.g., JSON parsing) is pushed to an async queue "to not block the response" — the user sees a response without the structured fields and must poll for the result, which is another full round trip.
Principle 4: Streaming beats waiting for the result
Why. The user's sense of "frozen" is driven by time to first token (TTFT), not end-to-end latency. For the same 5-second generation, streaming shows the first character at 200 ms and feels like "it's thinking," while non-streaming shows the complete result after 5 seconds and feels like "it's dead." This is the core of experience engineering — perceived time ≠ actual generation time.
How. Online services (chat, completion) default to streaming, using SSE (Server-Sent Events):
GET /v1/chat/completions
Content-Type: text/event-stream
data: {"choices":[{"delta":{"content":"Paged"}}]}
data: {"choices":[{"delta":{"content":"Attention"}}]}
data: {"choices":[{"delta":{"content":"is"}}]}
data: {"choices":[{"delta":{"content":"vLLM"}}]}
data: [DONE]No WebSocket (bidirectional, complex), no polling (a request per second is N extra requests). See the "streaming chat API" project in Portfolio Projects.
Counterexample. Counterexample one: the chat API is designed as "wait for the model to finish, return everything at once" — TTFT is 5 seconds and the user is gone. Counterexample two: polling simulates streaming (the frontend requests increments every 200 ms) — every poll is a full HTTP round trip, QPS multiplies by orders of magnitude, and the service gets hammered. Counterexample three: WebSocket for streaming — the bidirectional channel is never used, and complexity goes up. SSE is the standard answer for streaming output.
Principle 5: Check bottlenecks at all three layers — model, operator, system
Why. An inference service's bottleneck can sit at three layers: the model layer (too many FLOPs), the operator layer (poor kernel implementation), and the system layer (scheduling, batching, network). GPU utilization alone cannot tell you which — 100% GPU can be compute-bound (model layer) or memory-bound (operator layer). See The Roofline Model and Compute Analysis and GPU Architecture and Optimization.
How. Use the four-step diagnosis from Tuning and Performance Optimization:
Step 1: nvidia-smi — GPU utilization, memory utilization, temperature
Step 2: vLLM logs — KV cache usage, throughput, swapped
Step 3: Nsight Systems — operator time distribution
Step 4: Locate the bottleneck layer:
- Model layer → quantize, distill, prune
- Operator layer → FlashAttention, kernel fusion
- System layer → tune max_num_seqs, add PagedAttentionThe remedies are completely different per layer — blind tuning is straining at the wrong layer.
Counterexample. Counterexample one: GPU utilization is 100% but slow; the team assumes insufficient compute and adds GPUs — it was memory-bound, and adding GPUs doesn't fix memory bandwidth. Counterexample two: single-operator benchmarks look great (attention 5 ms) but end-to-end is slow — the bottleneck is the scheduling layer (kernel launches too frequent), and more operator optimization won't help. Counterexample three: only throughput is watched, never latency — p99 silently spikes, unnoticed. See Principle 9.
Principle 6: Look at activation outlier distribution before quantizing
Why. Quantization compresses continuous values into discrete buckets — models with a normal (bell-shaped) activation distribution quantize losslessly, while models with many activation outliers (like the "massive activations" of the LLaMA family) collapse. Quantizing without checking the activation distribution is gambling. This is the core trade-off of Model Quantization Fundamentals and the root motivation of Weight-Only Quantization and Mixed Precision — why is weight-only quantization more robust than weight+activation quantization in LLM scenarios? Because activation outliers are far more severe than weight outliers.
How. Run an activation distribution check before quantizing:
python
# Run a few dozen samples through transformers; statistics per layer
from transformers import AutoModel
model = AutoModel.from_pretrained("meta-llama/Llama-2-7b-chat-hf")
# Hook each layer's activations while running samples
activations = {}
def hook(name):
def fn(module, input, output):
activations[name] = output.detach()
return fn
for name, module in model.named_modules():
if "mlp.gate" in name or "attention" in name:
module.register_forward_hook(hook(name))
# After running samples, check each layer's max/mean ratio
for name, act in activations.items():
ratio = act.max() / act.mean()
print(f"{name}: max/mean = {ratio:.1f}") # > 100 means many outliersFor layers with ratio > 100, be careful with quantization — consider weight-only (keep activations at full precision), or activation-aware schemes like AWQ/GPTQ. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
Counterexample. Counterexample one: a team runs SmoothQuant (activation quantization) on LLaMA-2 and MMLU drops 3 points — LLaMA's activation outliers got crushed. Counterexample two: INT8 weight + INT8 activation on LLaMA runs slower than bf16 — dequantization overhead exceeds the memory savings (see "INT8 weights + FP16 activations compute slower" in Common Pitfalls and Anti-Patterns). Counterexample three: AWQ ships without a business evaluation — business quality silently degrades: MMLU barely moves, but the business prompt distribution drops 5 points.
Principle 7: Speculative decoding depends on batch size
Why. Speculative decoding gains depend heavily on batch size — at small batch, the large model's compute sits idle and every correct small-model guess is free; at large batch, the large model is already saturated, and the draft model only adds memory and compute. This is the core trade-off of Speculative Decoding and Medusa/EAGLE, and one of the most common anti-patterns.
How. The "golden range" of speculative decoding is batch ≤ 16 — beyond 32 is the break-even point, beyond 64 a net loss. In production:
| Scenario | Batch size | Speculative decoding | Rationale |
|---|---|---|---|
| Online chat | 1–16 | On | Gains are largest at small batch |
| Online batch | 32–64 | Off | Regresses at large batch |
| Offline batch | 256+ | Off | Pure net loss |
If your business has both "small-batch online" and "large-batch offline" traffic, split the traffic with a router — speculative-decoding instances take the small-batch traffic, regular instances take the large-batch traffic. See the "multi-model routing" project in Portfolio Projects.
Counterexample. Counterexample one: EAGLE enabled for a large-batch offline scenario — throughput regresses 30% — the large model is already compute-saturated, and the draft model just occupies memory. Counterexample two: online chat without speculative decoding — TTFT can't go below 50 ms — small-batch is the sweet spot; not enabling it wastes free performance. Counterexample three: unaware of the batch dependency, you set max_num_seqs to 64 with EAGLE on, then debug for a week without finding the cause — it is a design conflict at the root.
Principle 8: Gray-release at 5% traffic first
Why. New configurations (new quantization, new engine, new parameters) always have unexpected side effects — accuracy drift, long-tail latency, memory leaks, client compatibility. Full rollout = exposing everyone to the problem. A gray release contains the problem within 5% of traffic: small blast radius, fast rollback. This is standard practice in Model Serving and Orchestration and the canonical use of the multi-version capability in Triton Inference Server.
How. The standard gray-release flow:
Day 0: Deploy the new config to 1 instance (5% of instances)
- Traffic split 5:95 new:old by weight
- Watch QPS / p99 / error rate / business metrics for 24 hours
Day 1: No anomalies → expand to 20%
- Another 24 hours of observation
Day 2: No anomalies → expand to 50%
- From here a problem hurts more but stays controllable
Day 3: No anomalies → 100% full rollout
- Decommission the old configEvery step needs an automatic rollback mechanism: p99 beyond 1.5× SLA → auto-rollback; error rate > 1% → auto-rollback; business metrics drop 5% → auto-rollback. See "gray release and rollback" in Model Serving and Orchestration.
Counterexample. Counterexample one: full rollout of AWQ INT4, business quality collapses — MMLU unchanged, but the business prompt distribution drops 10 points. Rollback takes an hour; the business suffers in the meantime. Counterexample two: the new version is tested for latency only, not business metrics — latency looks fine, but the output format changed (INT4 shifted the logits), and client parsing fails. Counterexample three: no rollback plan prepared — when the problem surfaces, the team scrambles; the rollback takes longer than the problem itself.
Principle 9: Monitoring is more than QPS
Why. QPS is "traffic," not "health" — QPS can be normal while p99 silently spikes, KV cache hit rate drops, or GPU thermal throttling kicks in. These are all "looks fine, actually collapsing" signals. Production monitoring must be multi-dimensional, each with its own alert threshold.
How. The minimal monitoring set for an inference service:
| Metric | Meaning | Alert threshold |
|---|---|---|
| QPS | Traffic | Sudden change > 50% |
| p50 / p99 latency | Perceived quality | p99 > 1.5× SLA |
| GPU utilization | Compute usage | < 50% sustained for 5 minutes |
| GPU memory utilization | Memory usage | > 95% |
| GPU temperature | Cooling | > 85°C |
| KV cache hit rate | Prefix reuse | < 50% |
| Swapped request count | Memory swap-out | > 0 |
| Error rate | Anomalies | > 1% |
| Business metrics | End-to-end | Drop of 5% |
The three bold items are unique to LLM inference and most often ignored. See "vLLM logs under the microscope" in Tuning and Performance Optimization.
Counterexample. Counterexample one: only QPS and error rate monitored — p99 silently climbs from 200 ms to 800 ms, discovered only via user complaints. Counterexample two: GPU utilization monitored without temperature — thermal throttling silently costs 30% of performance, unnoticed until users ask "why is it slower?" Counterexample three: KV cache hit rate ignored — in a multi-turn scenario, prefix caching silently fails and TTFT doubles.
Principle 10: Failure drills as routine
Why. Failure is not "if" but "when." Production will meet: instance crashes, GPU failures, network partitions, full disks, dependency timeouts. Without drills, the failure finds you unprepared — you say "we have redundancy," and then discover failover was never configured; you say "we have monitoring," and then discover the alert thresholds were wrong. Netflix's Chaos Monkey thinking applies equally to inference services: break things on purpose and verify the system's resilience.
How. Drill at least four failure types:
| Failure type | Drill method | What to verify |
|---|---|---|
| Single-instance crash | kubectl delete pod — kill a random pod | Does traffic shift automatically? Does p99 stay within SLA? |
| GPU failure | nvidia-smi -r — reset the GPU | Does the instance restart automatically? Does the alert fire? |
| Network partition | iptables simulating packet loss | Does the client retry? Does a retry storm (cascading failure) form? |
| Upstream timeout | Simulate 5-second LLM responses | Is degradation triggered (route to the small model or return cache)? |
Drill monthly; write a postmortem each time. See "failure drills" in Model Serving and Orchestration.
Counterexample. Counterexample one: no drills — when a GPU fails, the instance restarts automatically but the KV cache is lost; every in-flight request fails and client retry storms cascade. Counterexample two: during a drill, you discover failover was misconfigured — killing a pod returns 503s to users instead of shifting traffic. Counterexample three: no degradation for upstream timeouts — a single LLM timeout stalls the entire request chain.
Principle 11: Quantization is a trade-off, not a free lunch
Why. Quantization is one of the most effective inference optimization techniques, but it always has a cost — accuracy loss, tuning complexity, engine compatibility. Every team that treated quantization as a silver bullet has been burned: INT4 shipped, business quality silently degrades; MMLU unchanged but JSON output formats break; quantization results inconsistent across engines. See "the trade-offs of quantization" in Model Quantization Fundamentals.
How. A quantization decision must answer four questions:
Q1: How tolerant is the business of accuracy loss?
High (chat, summarization) → AWQ INT4 is an option
Low (math, code) → stay bf16, INT8 at most
Q2: What is the model's activation outlier distribution?
Normal → weight+activation quantization is viable
Many outliers → weight-only quantization only
Q3: Has the business evaluation set been run before and after?
Run, < 1 point drop → ship
Not run → do not ship
Q4: Are the quantization costs accounted?
Gains: 1.5–2× throughput, 1/4 memory
Costs: accuracy loss, engine compatibility, debugging complexitySee Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
Counterexample. Counterexample one: INT4 shipped without a business evaluation — MMLU barely moves but the business prompt distribution drops 5 points. Counterexample two: activation quantization without checking activation distribution — LLaMA outliers crushed, output format changes. Counterexample three: assuming quantization always speeds things up — but with a small model + large batch, dequantization overhead exceeds the memory savings, and performance actually regresses.
Principle 12: End-to-end benchmarks, not single-operator benchmarks
Why. A single-operator benchmark (e.g., "attention in 5 ms") reflects engine implementation quality but cannot reflect user experience. What the user feels is end-to-end: HTTP request → tokenizer → prefill → decode → detokenizer → content filter → HTTP response. Attention is just one segment; the bottleneck can be entirely elsewhere. Making decisions with single-operator benchmarks is viewing a panorama through a microscope.
How. Deployment decisions must be based on end-to-end benchmarks:
python
# End-to-end latency (full HTTP request → complete response)
import time, requests
t0 = time.perf_counter()
resp = requests.post("http://llm-service/v1/chat/completions",
json={"messages": [...]})
t1 = time.perf_counter()
print(f"End-to-end: {(t1-t0)*1000:.0f} ms")Report both "end-to-end latency" and "in-engine latency" (measured by vLLM itself) — the former is user perception, the latter engine capability. See "what to measure" in Inference Benchmarking in Practice.
Counterexample. Counterexample one: the team optimizes the attention operator from 5 ms to 3 ms, but end-to-end latency doesn't move — the bottleneck is HTTP and tokenization. Counterexample two: single-operator benchmarks look great, but after launch, poor scheduling makes end-to-end twice as slow. Counterexample three: when switching engines, only in-engine benchmarks are compared, never end-to-end — the new engine is faster internally but slower in the HTTP layer, a net regression.
3. Tensions Between Principles
The 12 principles are not independent — they pull against each other and must be balanced per scenario:
Tension 1: Principle 1 (latency vs throughput) vs Principle 8 (gray release)
A latency-first config (small batch) may look fine during gray release because "traffic is small," surfacing problems only at full rollout. Response: during gray release, deliberately overdrive traffic (5% of instances carrying 20% of traffic) to stress-test.
Tension 2: Principle 4 (streaming) vs Principle 9 (monitoring)
Streaming latency distributions differ completely from non-streaming — you cannot reuse the same monitoring thresholds. Response: for streaming services, p99 must be chunk-level latency, not request-response latency. See "streaming test methods" in Inference Benchmarking in Practice.
Tension 3: Principle 6 (activation outliers) vs Principle 11 (quantization is a trade-off)
A model with many activation outliers goes weight-only, but weight-only accelerates less than weight+activation on some engines. Response: check business accuracy tolerance first, then pick the scheme — in accuracy-sensitive scenarios, keep weight-only even if slower.
Tension 4: Principle 7 (speculative decoding vs batch) vs Principle 1 (latency vs throughput)
Latency-first (small batch) happens to be the sweet spot of speculative decoding, while throughput-first (large batch) means turning it off. Response: the two don't actually conflict — latency-first picks small batch + speculative decoding, throughput-first picks large batch + no speculative decoding. They are two sides of the same decision.
4. Turning Principles into Team Norms
The 12 principles only bind when codified into team norms. Suggested structure:
markdown
# Team Deployment Standards
## 1. Pre-deployment checklist
- [ ] Business SLA defined (P1)
- [ ] Memory budget calculated (P2)
- [ ] Streaming / non-streaming decided (P4)
- [ ] Quantization scheme evaluated (P6, P11)
- [ ] Speculative decoding applicability evaluated (P7)
## 2. During-deployment norms
- [ ] Pre/post-processing in the same service as the model (P3)
- [ ] Three-layer bottleneck diagnosis completed (P5)
- [ ] End-to-end benchmark measured (P12)
## 3. Launch norms
- [ ] Gray-release process (P8)
- [ ] Monitoring and alert thresholds (P9)
- [ ] Failure drill plan (P10)
## 4. Operations norms
- [ ] Monthly failure drills
- [ ] Quarterly SLA review
- [ ] Postmortem for every incidentSee "team standards template" in Model Serving and Orchestration.
5. Further Reading
- Deploy an Inference Service from Scratch — end-to-end application of the principles
- Progressive Tutorial: Three Working Versions — verifying the principles one by one
- Inference Benchmarking in Practice — the methodology of Principle 12
- Tuning and Performance Optimization — the tools for Principles 5 and 7
- Common Pitfalls and Anti-Patterns — the counter-list of these principles
- Model Serving and Orchestration — the theory behind Principles 3, 8, 9, 10
- Batching and Request Scheduling — the theory behind Principles 1 and 7
- Latency, Throughput, and Concurrency — the theory behind Principles 1 and 4
- The GPU Memory Hierarchy and the Bandwidth Wall — the theory behind Principles 2 and 5
- The Roofline Model and Compute Analysis — the tool for Principle 5
- Model Quantization Fundamentals — the theory behind Principles 6 and 11
- Weight-Only Quantization and Mixed Precision — the deep dive of Principle 6
- GPU Architecture and Optimization — the tool for Principle 9
- Triton Inference Server — the implementation of Principle 8
- Speculative Decoding and Medusa/EAGLE — the deep dive of Principle 7
References
- Google: Rules of ML — the granddaddy of deployment design principles
- NVIDIA: Deploying LLMs in Production — Triton deployment guide
- vLLM: Production Best Practices — vLLM in production
- SRE Book: Service Level Objectives — how to define SLAs
- Netflix: Chaos Engineering — failure drill methodology
- LLM Serving: Continuous Batching — batch scheduling
- PagedAttention Paper — KV cache management
- LLM Quantization: AWQ — activation-aware quantization
- Speculative Decoding: EAGLE — the batch dependency of speculative decoding