Appearance
Common Pitfalls and Anti-Patterns
Lessons from inference deployment are always worth more than tricks — every pitfall collected on this page comes from a real project's painful moment.
LLM inference deployment projects rarely fail because the engine isn't advanced enough. They fail because a dumb mistake was made somewhere invisible: shipping quantized models without running the business evaluation, preprocessing inconsistent with training, single-batch latency mistaken for production-grade latency, KV cache capacity miscalculated into a day-one OOM, GPU utilization used as the sole performance metric... None of these errors appear in any engine's "quick start," yet each one decides whether a deployment ships or gets reworked.
This article organizes 15 high-frequency pitfalls into five layers: model / data / scheduling / monitoring / engineering. Each pitfall follows a uniform format — symptom / root cause / fix / related pages — and a tickable self-check list is attached at the end. Treat this page as a "pre-operative checklist": walk through it before every deployment and it will save you from nine out of ten rework cycles.
Establish the frame first
If you don't yet have a systematic deployment framework, read Deploy an Inference Service from Scratch and Deployment Design Principles first — load the full picture of "business SLA → memory budget → engine selection → tuning → launch" into your head. The "counter-list" in this article will then land much better.
1. Quantization and Accuracy Pitfalls
Quantization is the most effective lever for LLM inference optimization — and the easiest to botch. Its side effects don't show up on MMLU; they surface quietly on the business prompt distribution.
Pitfall 1: Accuracy not aligned before and after quantization (shipping without running any evaluation set)
The AWQ INT4 model ships because MMLU barely moved, and only after launch do you discover the output format is broken.
Symptom: MMLU drops less than 1 point — looks "essentially lossless." But business scenarios (code generation, JSON output, long-document summarization) degrade badly — JSON field formats break, code syntax errors appear, summaries miss key information. The most insidious case: users report "it got dumber" while MMLU before and after quantization is nearly identical.
Root cause: MMLU is a multiple-choice evaluation — it checks "which option has the highest probability." INT4 quantization shifts the logits of all tokens; the top-choice answer doesn't change, but the absolute probability distribution does — invisible in multiple choice, amplified in "generate a complete token sequence" tasks. Code and JSON, where format depends on every token, break entirely when one goes wrong.
Consequence: business quality quietly degrades after launch; the team assumes "the model itself isn't capable enough" and switches to a bigger model — the problem isn't solved, and new complexity is added. Worst case: users find the problem first, and the team reacts passively.
Fix: the business evaluation set must be run before and after quantization — not MMLU, but accuracy on your real prompt distribution, format validity, and human scoring. Minimum evaluation set:
python# Business evaluation script (pseudocode) prompts = load_business_prompts(n=200) # Sample of real business prompts for prompt in prompts: bf16_out = bf16_model.generate(prompt) awq_out = awq_model.generate(prompt) # Three metrics bleu = compute_bleu(bf16_out, awq_out) # Character similarity json_valid = check_json(awq_out) # Format validity rate human_score = human_eval(awq_out, prompt) # Human scoring (50 samples)Ship only if all three pass (bleu > 0.95, json_valid > 95%, human_score on par). See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
Pitfall 2: Preprocessing inconsistent with training (wrong tokenizer, inconsistent padding)
Training used the fast tokenizer + left padding; inference uses the slow tokenizer + right padding — output quality suddenly collapses.
Symptom: inference output is noticeably worse than during training — repetition, incoherence, answers missing the point. Inspection shows the tokenizer "looks identical" to training, but padding direction, special tokens, and max_length all differ.
Root cause: "preprocessing consistency between training and inference" is a multi-layer concept — tokenizer version consistency (fast vs slow behave differently), padding direction (left vs right affects generation), special token consistency (BOS/EOS added or not), max_length consistency, tokenizer config consistency (add_special_tokens=True/False). Any layer diverges, and the token sequence reaching the model differs.
Consequence: the model acts "dumber"; the team burns a week tracing it to a tokenizer config mismatch — maddening to locate, because "everything looks normal."
Fix: training and inference share the same preprocessing code (one function, one config file). Before launch, run the same prompt through both the training framework and the inference engine, and compare token id sequences for exact equality:
python# Training framework train_tokens = tokenizer.encode(prompt, return_tensors="pt") # Inference engine (vLLM also uses a tokenizer internally) vllm_tokens = tokenizer.encode(prompt, return_tensors="pt") assert torch.equal(train_tokens, vllm_tokens), "Token sequences do not match!"See "pre/post-processing in the same service as the model" in Deployment Design Principles.
Pitfall 3: INT8 weights + FP16 activations compute slower
Weights quantized to INT8 but activations in FP16: dequantization overhead exceeds the memory savings, and performance gets worse instead of better.
Symptom: INT8 weight-only quantization shipped, expecting a 1.5–2× speedup; throughput is basically flat or slightly worse — and nvidia-smi shows GPU utilization actually dropped.
Root cause: weight-only quantization's logic is "weight traffic halved, activations at full precision." But INT8 weights must be dequantized back to FP16/FP32 every step for the matmul — and that dequantization is extra operator overhead. At small batch (≤ 4) or small models (≤ 3B), dequantization cost > memory savings, so performance regresses.
Consequence: the team's quantization work delivers zero performance gain while losing accuracy — a double hit.
Fix: weight-only quantization only pays off at large model + medium batch:
Model size Batch INT8 weight-only gain ≤ 3B Any Negative (dequantization dominates) 7B+ 1–4 Slightly positive (10–20%) 7B+ 8–64 Positive (30–60%) 7B+ 128+ Strongly positive (near 2×) For small models at small batch, either skip quantization or go full weight+activation quantization (despite the bigger accuracy loss). See Weight-Only Quantization and Mixed Precision and The Roofline Model and Compute Analysis.
2. Scheduling and Concurrency Pitfalls
The scheduling layer is where performance differs the most — the same model on the same hardware can differ 10–30× between good and bad scheduling. These three pitfalls are the ones beginners hit most.
Pitfall 4: Single-batch latency mistaken for production-grade latency
A single request to vLLM shows TPOT 20 ms, and you promise the business "20 ms latency on one GPU."
Symptom: single-request testing shows a beautiful TPOT of 20 ms; you promise "10 tokens within 200 ms." After launch, the business pushes 32 concurrent requests — TPOT flies to 50 ms+, p99 to 200 ms+ — a massive SLA breach.
Root cause: single-request testing is the engine's best case — KV cache uncontended, scheduling unblocked, GPU compute idle. Production reality: 32 concurrent requests arrive together, the KV pool runs short and swaps, the scheduler queue backs up, and GPU compute gets divided. Single-request latency reflects the "engine capability ceiling," not the "production load floor."
Consequence: the business forms wrong latency expectations, the SLO is set wrong, and user experience collapses after launch. Worst case: SRE keeps chasing "why is production slower than testing," never realizing the test method was wrong.
Fix: always report single-request latency and N-concurrency latency together; ideally plot a "concurrency vs latency" curve:
Concurrency p50 p99 1 22 25 ← engine capability ceiling 4 24 30 16 28 50 ← recommended production ceiling 32 35 180 ← inflection point, degradation begins 64 50 1200 ← severe degradationProduction SLOs should be set from "before the inflection point." See Inference Benchmarking in Practice and Batching and Request Scheduling.
Pitfall 5: KV cache capacity misestimated (long conversation history blows memory)
Deployment sized the KV cache for "prompt 500 + output 256"; after launch, users chat for 20 turns and memory blows up.
Symptom: single-turn tests all pass. After launch, users run multi-turn conversations (each turn appends history); around turn 10, OOMs start, the KV pool bursts, and swapped counts spike.
Root cause: multi-turn prompt length accumulates — every turn adds the prior conversation history, and after 10 turns a prompt easily exceeds 4000 tokens. The deployment budget assumed "average prompt 500," while the production long tail can reach 8000+. KV cache memory = max_model_len × max_num_seqs × memory per token — underestimate any variable and it blows.
Consequence: everything works at launch, then collapses as users accumulate history — a "delayed failure" that is hardest to localize, because early data looks fine and the problem builds with usage time.
Fix: budget the KV cache by the business long-tail prompt length, with 50% headroom:
python# Business prompt length statistics prompt_lens = [len(tokenizer.encode(p)) for p in business_prompts] print("p50:", np.percentile(prompt_lens, 50)) print("p99:", np.percentile(prompt_lens, 99)) print("max:", max(prompt_lens)) # KV cache budget = p99 × max_num_seqs × memory per token × 1.5 (headroom)See "calculate the memory budget first" in Deployment Design Principles and "max_model_len tuning" in Tuning and Performance Optimization.
Pitfall 6: GPU utilization as the only metric (memory-bound shows 100% but slow)
nvidia-smi shows 100% GPU utilization; you conclude "compute is saturated" and keep adding batch with no gain.
Symptom: GPU utilization stays at 100%; the team decides "compute is maxed out" and adds GPUs. Throughput doesn't move — because the bottleneck is memory bandwidth, not compute.
Root cause: GPU utilization does not distinguish compute-bound from memory-bound — both show 100%. Under memory-boundedness, compute is actually idle (waiting for data to stream from VRAM to SRAM), but nvidia-smi still reports 100%. Using GPU utilization as the sole metric yields the false conclusion "compute saturated."
Consequence: real money spent on GPUs that doesn't help — the bottleneck is memory bandwidth, and more GPUs don't add bandwidth. Or the team tunes in the wrong direction (adding batch), making the memory-boundedness worse.
Fix: watch GPU utilization and memory utilization together:
GPU utilization Memory utilization Bottleneck layer Response > 90% < 50% Compute-bound Add GPUs, quantize (cut FLOPs) < 50% > 90% Memory-bound Weight-only quantization (cut traffic) > 90% > 90% Both saturated Need more GPUs and more bandwidth < 50% < 50% Poor scheduling Tune max_num_seqs, enable continuous batching See "Step 1: nvidia-smi macro view" in Tuning and Performance Optimization and The Roofline Model and Compute Analysis.
Pitfall 7: Large batch for online inference (throughput up, p19 shattered)
To "save cost," max_num_seqs goes from 16 to 64 — throughput triples, but p99 spikes 5–10×.
- Symptom: throughput numbers look great (QPS doubled), but users report "occasional freezes." Monitoring shows p50 unchanged, but p99 jumped from 200 ms to 1200 ms — a small minority of requests severely slowed.
- Root cause: large batch saturates GPU compute (throughput up), but the scheduler queue grows longer (every request waits longer, p99 up). The throughput–p99 trade-off is nonlinear — p99 grows far faster than throughput. This is the core trade-off of Batching and Request Scheduling.
- Consequence: p99-sensitive businesses (chat, completion) see "a few users' experience collapse" — and those few users complain and churn. Overall QPS up 30% while 5% of users leave is a bad deal.
- Fix: set
max_num_seqsbefore the "concurrency–p99 inflection point" — see "max_num_seqs tuning" in Tuning and Performance Optimization. For online serving, prefer losing some throughput to protecting p99 — the canonical application of "latency-first vs throughput-first: pick one first" in Deployment Design Principles.
Pitfall 8: Speculative decoding at large batch (net loss)
To "go faster," EAGLE is enabled in all scenarios — and the large-batch offline scenario regresses 30%.
- Symptom: EAGLE speculative decoding enabled; single-request latency indeed halves (small batch benefits). But in the 32-concurrency stress test, throughput regresses 30% versus EAGLE off.
- Root cause: speculative decoding gains depend heavily on batch size. At small batch, the large model's compute is idle and every correct small-model guess is free; at large batch, the large model is saturated and the draft model only adds memory and compute. Thirty-two concurrency is the break-even point; 64+ is a net loss.
- Consequence: the team debugs for a week without finding the cause — assuming an EAGLE configuration problem, they tune
num_speculative_tokensand make it worse. The real issue is a scenario mismatch. - Fix: enable speculative decoding only at small batch (online chat, code completion, batch ≤ 16). Disable it for large-batch scenarios (offline batch, high-concurrency online). If the business has both "small-batch online" and "large-batch offline" traffic, split it with a router — see the "multi-model routing" project in Portfolio Projects. See Speculative Decoding and Medusa/EAGLE and "speculative decoding depends on batch size" in Deployment Design Principles.
3. Monitoring and Operations Pitfalls
Monitoring-layer pitfalls are the stealthiest — everything "looks normal" until the user complaints arrive.
Pitfall 9: No GPU temperature monitoring (thermal throttling costs 30%)
Only QPS and latency monitored; no GPU temperature. A hot summer day throttles the GPUs 30%, and the team burns a week on it.
Symptom: QPS and error rate normal, but p99 latency suddenly climbs from 200 ms to 280 ms. Logs look fine, vLLM config untouched, model version unchanged. Until someone visits the data center and finds a dead GPU fan and 90°C temperatures.
Root cause: beyond a temperature threshold (usually 85°C), GPUs throttle automatically for protection — SM clocks and memory clocks drop, and performance falls 20–40%. This is hardware self-protection: no error, no warning — performance just quietly degrades.
Consequence: long diagnosis ("logs all normal" is the most misleading signal) and degraded user experience. Worst case: the data center's cooling was always insufficient, GPUs were always throttled, and performance never met spec.
Fix: monitor GPU temperature — alert at 80°C, emergency alert at 85°C:
bash# Monitoring command nvidia-smi --query-gpu=temperature.gpu,clocks.sm,clocks.mem,power.draw --format=csv -l 10 # Prometheus exporter exposes metrics # Metric names: nvidia_gpu_temperature_celsius, nvidia_gpu_clocks_sm_mhzAbove 85°C, alert automatically; above 90°C, automatically shed load (lower
max_num_seqs) or migrate the instance. See the "performance tuning checklist" in Tuning and Performance Optimization and "monitoring is more than QPS" in Deployment Design Principles.
Pitfall 10: ONNX export never verified (numeric drift creeps in)
A PyTorch model is exported to ONNX for ONNX Runtime deployment; results are wrong, and a week passes before anyone notices.
Symptom: the ONNX Runtime deployment ships; running the evaluation set reveals a 1–3 point accuracy drop. Initially blamed on ONNX Runtime implementation differences; a week of tuning later, it turns out to be numeric drift introduced during export.
Root cause: the PyTorch → ONNX export involves operator mapping (some PyTorch operators have no ONNX equivalent), precision conversion (some operators overflow in fp32 → fp16), and dynamic-shape handling (some dynamic dims get fixed during export). Every stage can introduce numeric drift — small per operator, significant in aggregate.
Consequence: weeks pass before anyone asks "why do ONNX results differ so much from the original model" — business quality has already silently degraded. Rolling back to the original model also takes time; the business suffers meanwhile.
Fix: ONNX exports must be verified — run the same inputs through PyTorch and ONNX Runtime and compare outputs:
python# PyTorch run pt_out = pt_model(input).detach().cpu().numpy() # ONNX Runtime run import onnxruntime as ort sess = ort.InferenceSession("model.onnx") onnx_out = sess.run(None, {"input": input.numpy()})[0] # Compare (tolerance to 3 decimal places) np.testing.assert_allclose(pt_out, onnx_out, rtol=1e-3, atol=1e-3)When drift exceeds the threshold, locate the responsible operator and replace it individually or keep PyTorch for that part. See ONNX Runtime: Cross-Platform and Computation Graph Optimization.
Pitfall 11: Engine upgrade without regression testing
vLLM upgraded from 0.5.x to 0.6.x; performance regresses 20%; it ships without regression testing.
Symptom: after the engine upgrade, benchmark numbers look normal (even slightly better), but production p99 spikes. Rolling back fixes it — the regression came from the upgrade.
Root cause: an engine upgrade can change default parameters (e.g.,
enable_chunked_prefillflipping from false to true), kernel selection (the new version may pick a different attention kernel), and quantization implementation (the INT8 compute path changed). These changes don't show on standard benchmarks but do show on your specific business prompt distribution.Consequence: business degrades after the upgrade; the team scrambles to roll back. Worst case: rollback is no longer possible (the new version's features are now depended upon), leaving emergency patching.
Fix: engine upgrades go through gray release + full regression testing:
markdown## Engine upgrade regression checklist 1. Run the complete business evaluation set in a test environment 2. Run the full benchmark suite from [Inference Benchmarking in Practice](/practice/benchmarking) 3. Compare new vs old: throughput, p50/p99, memory, KV hit rate 4. Human side-by-side comparison on sampled business prompts 5. Gray-release at 5% traffic, observe 24 hours 6. No anomalies → 20%, another 24 hours 7. No anomalies → full rolloutSee "gray-release at 5% traffic first" in Deployment Design Principles.
Pitfall 12: CUDA / cuDNN version mismatch silently degrading performance
torch 2.3 built for CUDA 12.1, but production has CUDA 12.4 installed; performance drops 15% and nobody notices.
Symptom: local testing is perfect; the production deployment is 15% slower. Identical config, identical weights, identical parameters — just slower.
Root cause: GPU software like PyTorch / vLLM / TensorRT has strict dependencies on CUDA and cuDNN versions — same major version but mismatched minor versions can demote some kernels (using the slow implementation instead of the fast one). This degradation throws no error and no warning — it just quietly slows down.
Consequence: long diagnosis — everything "looks right," yet performance is off. The team suspects the model, parameters, hardware, and burns days before discovering a CUDA minor-version mismatch.
Fix: lock the entire software stack with Docker images:
dockerfileFROM nvcr.io/nvidia/pytorch:24.07-py3 # This image pins the torch 2.4 + CUDA 12.4 + cuDNN 9.0 combination # As long as the deployment machine has a sufficient NVIDIA driver, # the entire software stack stays consistentOr lock it with conda:
bashconda install pytorch=2.3.1 pytorch-cuda=12.1 cudatoolkit=12.1 cudnn=9.0 -c pytorch -c nvidiaSee "end-to-end benchmarks, not single-operator benchmarks" in Deployment Design Principles — only end-to-end benchmarks catch this class of problem.
Pitfall 13: Multiple models sharing a GPU without MPS / cgroup
Three services — vLLM + embedder + reranker — share one GPU with no MPS; they fight over memory and compute, and performance swings wildly.
Symptom: each service performs fine alone; together, performance swings violently — vLLM's p99 fluctuates between good and bad, and the embedder occasionally times out.
Root cause: by default, multiple CUDA processes on one GPU preempt each other's memory and compute — vLLM fills the memory and the embedder OOMs at startup; vLLM runs attention and the embedder's kernels get suspended. With no isolation mechanism, every process thinks it owns the GPU.
Consequence: unstable performance, sporadic OOMs, unpredictable scheduling. Worst case: one service's memory leak fills the entire GPU and all three services crash.
Fix: use MPS (Multi-Process Service) or cgroup + memory quotas:
bash# Option 1: MPS (NVIDIA recommended) nvidia-cuda-mps-control -d # Then run all three services through MPS — shared compute with # fewer context switches # Option 2: cgroup + per-process GPU memory quotas # vLLM: gpu_memory_utilization=0.6 # embedder: 0.2 # reranker: 0.2 # Sum ≤ 1.0 with explicit headroomSee GPU Architecture and Optimization and "calculate the memory budget first" in Deployment Design Principles.
4. Protocol and Design Pitfalls
Pitfall 14: Streaming responses done with polling instead of SSE
The chat API is designed as "wait for the model, return everything at once," and the frontend simulates streaming by polling for increments every 200 ms.
Symptom: QPS inflates by an order of magnitude out of nowhere — each user polls 5 times per second; 32 concurrent users mean 160 QPS, and the service gets hammered. One generation request is effectively amplified into 5 HTTP requests.
Root cause: the backend never designed a streaming interface; the frontend "fakes streaming" with polling — every poll is a full HTTP request, far costlier than real streaming.
Consequence: the service gets hammered, QPS explodes, and ops costs double. Worst case: user experience is worse than "wait for the result" — a long polling interval feels frozen, a short one melts the service.
Fix: use SSE (Server-Sent Events) for streaming output:
pythonfrom fastapi import FastAPI from sse_starlette.sse import EventSourceResponse app = FastAPI() @app.get("/chat") async def chat(prompt: str): async def event_generator(): async for chunk in vllm_stream_client.generate(prompt): yield {"data": chunk.text} yield {"data": "[DONE]"} return EventSourceResponse(event_generator())See "streaming beats waiting for the result" in Deployment Design Principles and the "streaming chat API" project in Portfolio Projects.
Pitfall 15: Context length beyond the model's native support, faked with sliding window — accuracy collapses
The model natively supports 4K context, but the business needs 32K — implemented with a sliding window, and long-document accuracy collapses.
Symptom: sliding window handles 32K long documents; the model "forgets" key information from earlier — long-document summaries miss key sections, long conversations forget earlier commitments.
Root cause: a sliding window's essence is "only look at the most recent N tokens" — everything beyond the window is discarded outright. The model appears to "read a 32K document," but is actually seeing only the last 4K; the first 28K never entered the KV cache at all. This is not optimization — it is pretending to support long context.
Consequence: users report "the model got dumber" — the team assumes insufficient capability and swaps in a bigger model; the problem persists. The real issue is the wrong context-handling method.
Fix: either use a model with native long-context support (e.g., Qwen2.5-7B-Instruct with native 32K, Yarn-7B-128K), or use real long-context methods:
- Chunked prefill: long prompts enter the KV cache in chunks without losing earlier content (native in vLLM 0.5+)
- RoPE scaling: position interpolation extends a 4K model to 32K (slight accuracy drop, far better than sliding window)
- RAG: don't stuff the whole document into the model — use a retrieval + summarization hybrid
See "end-to-end benchmarks, not single-operator benchmarks" in Deployment Design Principles — a sliding window "looks like 32K support" on single-operator benchmarks while end-to-end quality collapses.
5. Self-Check Checklist
The 15 pitfalls condensed into a tickable pre-deployment checklist:
markdown
## Pre-Deployment Self-Check Checklist
### Quantization and accuracy
- [ ] Business evaluation set run before and after quantization (not just MMLU) — Pitfall 1
- [ ] Tokenizer fully consistent with training (fast/slow, padding direction) — Pitfall 2
- [ ] Weight-only INT8 used at large model + medium batch (not small model + small batch) — Pitfall 3
### Scheduling and concurrency
- [ ] Single-batch latency + multi-concurrency latency reported together; inflection found on the curve — Pitfall 4
- [ ] KV cache sized by business long-tail prompt length with 50% headroom — Pitfall 5
- [ ] GPU utilization and memory utilization watched together; compute/memory-bound distinguished — Pitfall 6
- [ ] max_num_seqs before the p99 inflection point; prefer losing throughput over losing p99 — Pitfall 7
- [ ] Speculative decoding only at small batch; off at large batch — Pitfall 8
### Monitoring and operations
- [ ] GPU temperature monitored (alert at 85°C, emergency at 90°C) — Pitfall 9
- [ ] ONNX exports verified for numeric consistency — Pitfall 10
- [ ] Engine upgrades go through gray release + full regression testing — Pitfall 11
- [ ] Entire CUDA/cuDNN software stack locked with Docker — Pitfall 12
- [ ] Multiple models on one GPU use MPS or memory quotas — Pitfall 13
### Protocol and design
- [ ] Streaming output uses SSE, not polling — Pitfall 14
- [ ] Long context via native support or chunked prefill, not sliding window — Pitfall 15How to use this checklist
- Before every new deployment: walk through it, confirming every item has been checked
- Before every upgrade: walk through it, confirming the upgrade introduces no new pitfall
- After every incident: walk through it, identifying which pitfall wasn't guarded
- Quarterly review: tally which pitfalls the team has hit, and update the checklist
6. Further Reading
- Deploy an Inference Service from Scratch — end-to-end deployment flow
- Progressive Tutorial: Three Working Versions — one optimization per version, compared
- Inference Benchmarking in Practice — the methodology behind Pitfalls 4, 5, 6
- Tuning and Performance Optimization — the tools behind Pitfalls 5, 6, 7, 9
- Deployment Design Principles — the inverse principles of these pitfalls
- Model Quantization Fundamentals — the theory behind Pitfalls 1, 3
- Weight-Only Quantization and Mixed Precision — the deep dive of Pitfall 3
- Batching and Request Scheduling — the theory behind Pitfalls 4, 7, 8
- Model Serving and Orchestration — the theory behind Pitfalls 11, 13, 14
- The Roofline Model and Compute Analysis — the tool behind Pitfalls 6, 8
- The GPU Memory Hierarchy and the Bandwidth Wall — the theory behind Pitfalls 5, 6
- GPU Architecture and Optimization — the tool behind Pitfalls 9, 13
- Kernel Fusion and Custom Kernels — the tool behind Pitfall 10
- Computation Graph Optimization — the deep dive of Pitfall 10
- ONNX Runtime: Cross-Platform — the engine behind Pitfall 10
- vLLM and PagedAttention — the engine behind Pitfalls 5, 7, 11
- Speculative Decoding and Medusa/EAGLE — the deep dive of Pitfall 8
- Latency, Throughput, and Concurrency — the theory behind Pitfalls 4, 7
References
- Google: Rules of ML — the granddaddy of training/serving skew and pitfall checklists
- vLLM: Production Best Practices — vLLM in production
- NVIDIA: Deploying LLMs in Production — Triton deployment guide
- ONNX Runtime: Model Verification — ONNX numeric verification methods
- NVIDIA MPS Documentation — multi-process service
- NVIDIA Nsight Systems — profiling tool
- SSE Specification — the Server-Sent Events protocol
- Continuous Batching: Orca Paper — the continuous batching paper
- PagedAttention Paper — KV cache management
- EAGLE-3: Speculative Decoding — the batch dependency of speculative decoding