Skip to content

Common Pitfalls and Anti-Patterns

At a glance A field guide to deployment failures: train/serve preprocessing skew, undetected quantization collapse, GPU memory leaks, launching with no monitoring, big-bang releases... every pitfall gets a symptom → cause → fix, all hard-won lessons, plus a 10-point pre-deployment checklist.

The classic pitfalls of model deployment form a small set of accident patterns whose symptoms look wildly varied but whose root causes keep repeating. Why does this deserve a dedicated page? Because deployment failures are remarkably predictable: 80% of the holes teams fall into are the same ten-odd ones, and each one takes hours to days to go from "production is broken" to "root cause found" — knowing these patterns in advance is like carrying a pre-loaded list of suspects into every investigation. This page collects the ten most classic pitfalls and anti-patterns in a symptom → cause → fix structure, cross-references the related page for each, and closes with a 10-point pre-deployment checklist. For the complete hands-on process, see Deploy a Model from Scratch.

How to Read This Page

Match the symptom first, then check the cause. When production breaks, jump to the matching entry and follow the investigation path in its "fix." Reading it once and then getting burned once beats memorizing it ten times.

Pitfall 1: Train/Serve Preprocessing Skew ​

  • Symptom: offline validation AUC is 0.92; after launch, accuracy craters by 10-20 points; code and environment both check out "identical."
  • Cause: training and serving preprocess inputs differently — three usual suspects: different normalization parameters (training uses mean=[0.5,0.5,0.5], serving hand-writes x/255 and forgets to subtract the mean); channel order (training RGB, serving BGR); different tokenizer versions (training on an old tokenizers release, serving on a new one). The model's assumptions about the input distribution get silently broken.
  • Fix:
    1. Write preprocessing exactly once: define it in preprocess.py, and have training and serving import the same function (the approach used in Deploy a Model from Scratch);
    2. Ship normalization parameters with the model: write them into meta.json at export time and have the inference class read them, eliminating "two hard-coded copies";
    3. Run an end-to-end consistency test before launch: the same image through the training pipeline vs the serving endpoint, asserted with np.allclose(atol=1e-4).

Why "Identical-Looking" Code Still Differs

The sneakiest variant is order of operations: training normalizes then resizes, serving resizes then normalizes — the numbers are completely different. Compare with numeric assertions, not your eyes.

Pitfall 2: Numerical Precision Gaps: FP16 Engines vs FP32 Training ​

  • Symptom: after switching to TensorRT/Triton for speed, a few requests return wrong results or bizarre scores; overall accuracy dips slightly but occasionally goes "off the rails."
  • Cause: inference engines default to FP16 (half precision) for speed, while training runs FP32. FP16's dynamic range is much narrower than FP32's (max ~65504), so models with wide normalized distributions or volatile numerics (especially RNNs and the logits just before an attention softmax) can overflow or lose precision.
  • Fix:
    1. Compare engine outputs against the original framework's outputs numerically, with a defined tolerance (FP16 typically allows differences on the order of 1e-2; differences reaching 1e0 mean something is wrong);
    2. Keep critical layers (like the logits before softmax) in FP32;
    3. Re-run the full offline evaluation whenever you switch engines — FP16 precision loss never throws an error; it just quietly degrades. For formats and engines, see Model Formats and Conversion.

Pitfall 3: Quantization Accuracy Collapse Left Unverified ​

  • Symptom: the model got "2x faster," but online conversion dropped 15% — and only afterwards does anyone notice the INT8 model was never run against the validation set.
  • Cause: the quantization pipeline ran; the validation step didn't — either the calibration data was wrong (random noise / a single class), or layers that are inherently sensitive got quantized anyway. INT8 squeezes activations down to 256 levels; get the distribution wrong and the error amplifies into a systematic bias.
  • Fix:
    1. After quantizing, always do three things: confirm the size (≈1/4 of FP32), compare validation-set accuracy (a drop > 5% means suspect the calibration data first), and measure latency for real;
    2. Calibrate with 200-1000 samples from the real distribution, covering all classes;
    3. Bake "post-quantization evaluation" into a script that runs in CI (the full workflow is in Model Optimization in Practice). For the theory, see Quantization.

Pitfall 4: GPU Memory Leaks and OOM ​

  • Symptom: after three days the service starts throwing OOM / out-of-memory errors; a restart fixes it, it relapses days later; P99 creeps upward over time.
  • Cause: four common sources — (1) long-lived connections / request objects never released (request bodies and image bytes still referenced); (2) unbounded engine cache growth (TensorRT workspaces, ONNX Runtime arena misconfigured); (3) a new session per request (constructing an InferenceSession for every request); (4) batching queues piling up without bound.
  • Fix:
    1. One global session: load the model once and share it across all requests (ONNX Runtime sessions are thread-safe — this is exactly what Deploy a Model from Scratch does);
    2. Load-test for ≥ 10 minutes while watching the memory curve (memory that never comes back down is a leak signal — see Load Testing and Capacity Planning);
    3. Set a container memory limit plus automatic OOM restart as the safety net, but learn to tell "leak" from "normal growth";
    4. Use tracemalloc / nvidia-smi to find the holder.

Pitfall 5: Worker Count and Duplicate Model Loads ​

  • Symptom: a container with 4 workers uses 4x the expected GPU/CPU memory; a 16GB card only fits 2 workers; slow model loads stretch startup time.
  • Cause: process-based workers (e.g. uvicorn --workers 4) each load their own copy of the model. With a 3GB model, 4 workers means 12GB. Beginners routinely assume workers share the model's memory. They don't.
  • Fix:
    1. Do the math first: worker count × model size ≤ 70% of memory/VRAM;
    2. Small model, low QPS → workers=2 or even 1 + async is plenty;
    3. Large model → use a shared memory / single-process multi-threaded architecture, or separate model loading from the request processes (e.g. Triton's concurrent model instances are managed by the engine — see NVIDIA Triton Multi-Model Serving);
    4. Warm up the model at startup (run 1-2 dummy requests after loading), or the first real request's latency will explode.

Pitfall 6: Flying Blind: Shipping Without Metrics ​

  • Symptom: launch with zero instrumentation; two weeks later the business team says "quality seems off," and you have no data to answer "since when, by how much, and which traffic is worst."
  • Cause: only functional verification happened before launch — no metrics instrumented, no monitoring hooked up. Model systems also have a special amplifier: data drift — the input distribution shifts, quality degrades accordingly, and without monitoring you never find out (Monitoring and Observability explains why drift must be watched).
  • Fix:
    1. Launch with the golden signals: RPS, P99, error rate, saturation (GPU utilization / queue depth);
    2. The acceptance bar for "metrics are live" is "up == 1 and someone is looking", not "the metrics endpoint responds";
    3. Drift detection: monitor input distributions (mean/variance/class frequencies) with threshold-based alerts;
    4. For complete working configs, see Observability in Practice.

Pitfall 7: Big-Bang Releases with No Canary ​

  • Symptom: the day after a full-volume launch, support tickets explode; when you try to roll back, there's no prepared command and the old image has been overwritten.
  • Cause: treating "offline validation passed" as sufficient for a safe launch. Model quality depends on the live data distribution; no offline set, however large, is more than a sample. A big-bang release outsources "validation" to real users.
  • Fix: canary by default: 5%→20%→50%→100%, observed stage by stage, with automatic rollback conditions configured and rehearsed ahead of time. Keep three-tier rollback (model / features / code) with version pairing. The full process is in Release Strategies: Canaries and Rollbacks.

Pitfall 8: Treating Online Inference Like Batch (or Vice Versa) ​

  • Symptom (in both directions):
    • Online direction: using a batch prediction endpoint one item at a time — every request ships a batch=256 input, the GPU computes 256 rows to return 1, and latency and cost explode;
    • Batch direction: pushing a million offline records one by one through the online service; three days later the service is down.
  • Cause: confusing two workload shapes. Online inference is low-latency, high-concurrency, one response at a time; batch is high-throughput, queueable, completed as a whole. Their resource strategies, timeout settings, and scheduling are entirely different.
  • Fix:
    1. Merge requests inside the online service with dynamic batching (Triton's dynamic batching is the engine-level solution — see the Triton case study), rather than asking callers to send big batches;
    2. Route large offline jobs through a batch pipeline (Argo/Celery/cloud batch), isolated from the online cluster — see the batch pipeline case study;
    3. Give the online service timeouts and rate limits so batch traffic can't drag down online responses (see the capacity model in Load Testing and Capacity Planning).

Pitfall 9: Cache Pollution and Consistency ​

  • Symptom: during A/B or canary testing, the quality comparison between versions is skewed; users report "the result changed after a refresh and now contradicts itself."
  • Cause: cache keys don't include the model version. When old and new models share a key, whoever requests first decides what later requests see — the "quality gap" during a canary may be nothing but a cache-hit difference, not a model difference.
  • Fix:
    1. Put model_version and feature_version into cache keys;
    2. Isolate volatile caches during the rollout (bypass the cache for the new version, or use a separate cache prefix);
    3. Be explicit about what's cached and for how long: result caches, feature caches, and tokenizer caches each need their own expiry policy;
    4. For the full discussion of consistency, see state management in Model Serving.

Pitfall 10: Load Test Numbers That Don't Match Production ​

  • Symptom: the load test says 1000 QPS at P99 50ms, but on day one real traffic of 300 QPS produces P99 500ms; or someone asks "where did these load test numbers come from?" and nobody can answer.
  • Cause: three classics — (1) insufficient warm-up: hammering the service right after startup while JIT/caches are still cold, producing numbers that are too low or too high; (2) inflated by cache: test requests all hit the cache (repeated parameters) while real traffic hits it zero times; (3) the load generator is the bottleneck: the test machines saturate first, so you measured the generator, not the service.
  • Fix:
    1. Warm up 30-60 seconds before recording numbers; run ≥ 10 minutes for steady-state scenarios;
    2. Randomize request parameters and separately measure the true "zero cache hits" path;
    3. Record the environment four-pack (load generator specs, worker count, concurrency, duration) so results are reproducible — the standard is in Load Testing and Capacity Planning.

The 10-Point Pre-Deployment Checklist ​

  • [ ] (1) Preprocessing defined in one place, normalization parameters shipped with the model, consistency test passing;
  • [ ] (2) Engine outputs numerically compared against the original framework, within expected tolerance;
  • [ ] (3) Post-quantization evaluated on the validation set, drop within the business threshold;
  • [ ] (4) Memory/VRAM curves flat over a 10-minute load test, no upward trend;
  • [ ] (5) Worker count × model size budgeted, and model warm-up at startup in place;
  • [ ] (6) Golden signals instrumented, up == 1, someone watching the dashboards;
  • [ ] (7) Releases go through a canary, automatic rollback conditions configured and rehearsed;
  • [ ] (8) Online/batch workloads separated, online service has timeouts and rate limits;
  • [ ] (9) Cache keys include versions, caches isolated during rollouts;
  • [ ] (10) Load test report documents environment, warm-up time, and parameter randomization — conclusions are trustworthy.

Checklist ​

  • [ ] You can name the three pitfalls your own service is most likely to hit, and the plan for each;
  • [ ] Incident reviews are recorded as symptom → cause → fix and archived in the team wiki;
  • [ ] The related page for each pitfall is bookmarked and read; any fix is findable within 10 minutes when needed;
  • [ ] The self-check list is wired into the release process (PR template / release ticket checkboxes).

Further Reading ​

References ​