Appearance
A model rollout is the process of shifting traffic from 0 to 100% onto a new model version, and the core goals are: quality gets validated in production, and rollback takes seconds when something breaks. Why is a model release so much more dangerous than a code release? Because when code releases fail, they crash; when model releases fail, they quietly get worse: the new model throws no exceptions — recommendations just get dumber, false positives creep up, users feel "it's not as good as it used to be," and you may not even have an alert for it. Add data drift on top, and even the best offline metrics can't save you from a production disaster. This page covers how to choose among the three mainstream strategies (canary, blue/green, A/B), the complete canary process, a three-tier rollback mechanism, and a pre-release checklist. For the theory, see The MLOps Pipeline; for a full gateway-based implementation, see Model Gateway and Canary Releases.
The Core Mindset for Releases
A release is not "flipping traffic" — it's a controlled experiment. Every release answers one question: is the new version better than the old one? The answer has to come from production metrics, not from the releaser's confidence.
1. Why Model Releases Are Riskier
| Dimension | Code release | Model release |
|---|---|---|
| Failure mode | Crashes, exceptions, 5xx | Quality degrades; the system runs "normally" but outputs get worse |
| Detection | Error rate, availability | Quality metrics (CTR / accuracy / conversion), with messy definitions |
| Rollback trigger | An error-rate alert is enough | Quality metrics degrade — but you first have to define "degrade" |
| Environment coupling | Code is tightly coupled to the environment | Also coupled to the data distribution: the same model can be fine today and bad tomorrow |
Data drift makes it worse: a new model performs well right after launch, then a week later the input distribution shifts and quality slides — what you shipped is not "the correct version," it's "the version that was correct under the current data distribution." That's why model releases must ship with monitoring and rollback from day one (for monitoring in practice, see Observability in Practice).
2. Comparing the Three Strategies: Canary, Blue/Green, A/B
| Strategy | How it works | Traffic split | Rollback speed | Quality validation | Best for |
|---|---|---|---|---|---|
| Canary | New version takes gradually increasing traffic, can retreat at any time | 5%→20%→50%→100% | Seconds (shift traffic back) | System + quality metrics observed at each stage | The default first choice |
| Blue/Green | Two full environments, switched as a whole | One switch 0%↔100% | Fast (switch back to the old environment) | No same-traffic comparison before the switch | Major infrastructure changes requiring a full cutover |
| A/B test | Long-running parallel split with statistical comparison | Fixed split (e.g. 50/50) | Depends on the validation period | Strict statistical significance | Validating quality, deciding algorithm iterations |
How to choose: canary by default (control, rollback, and validation in one package); blue/green for "full cutovers a canary can't handle" (database migrations, major dependency upgrades); A/B for "questions that need a rigorous answer on quality," usually combined with a canary — the canary keeps the release safe, and the A/B delivers the verdict on quality.
3. The Canary Release Process in Detail
3.1 Four Traffic Stages and What to Watch at Each One
The typical rhythm is 5% → 20% → 50% → 100%, pausing 10-30 minutes at each stage (for business-quality metrics, watch for at least one full business cycle).
| Stage | Traffic | What to watch | Pass criteria | On failure |
|---|---|---|---|---|
| 1 | 5% | System metrics: error rate, P99, GPU memory | Error rate ≤ old version, P99 not degraded | Roll back immediately; investigate environment differences |
| 2 | 20% | System + basic quality (accuracy / empty-response rate) | No significant drop in quality metrics | Roll back, or drop to 5% and retune |
| 3 | 50% | Business quality metrics (CTR / conversion / revenue) | On par with or better than the old version | Roll back and organize an analysis |
| 4 | 100% | Keep watching for 24-48 hours after full rollout | No regression, error budget not burning faster | The rollback plan takes effect immediately |
Key principle: each stage gets exactly one "no degradation" sign-off, and you never skip stages. Skipping means that when something breaks, your search space is two confounded variables: increased traffic plus a model change.
3.2 Automatic Rollback Conditions
Write rollback conditions as executable rules (enforced at the gateway/orchestration layer):
text
Automatic rollback triggers (any one of these rolls back):
- Error rate > 1.5x the old version for 5 minutes
- P99 latency > SLO threshold for 10 minutes
- A quality metric (e.g. conversion rate) down > 3% vs the old version for one business cycle
- Resource alarms (OOM, GPU memory exhausted)
Manual rollback triggers (anyone can invoke):
- Business-side reports of odd behavior, rising support tickets
- Any metric anomaly you can't explain: roll back first, investigate after3.3 What to Monitor During the Canary
A canary isn't "watching the dashboards" — it's watching a set of metrics designed for comparison: same metrics, both versions, one consistent methodology.
- System metrics: error rate, P99, and QPS for old and new versions (split by the
versionlabel — see the labeling standards in Observability in Practice); - Quality metrics: conversion rate, CTR, accuracy — the old version must keep serving comparable traffic the whole time, otherwise there's no control group;
- Drift signals: input distribution shifts — see Monitoring and Observability.
4. Blue/Green Deployment: The Big Switch and Its Price
Blue/green maintains two complete environments (green = current, blue = new). Once validation passes, incoming traffic switches to blue as a whole, then blue and green swap roles.
| Advantages | Costs |
|---|---|
| Switching is nearly instant (seconds); rolling back is just another switch | Double the resource cost (two full sets of GPUs/instances running all the time) |
| The new environment can be validated thoroughly beforehand | Config and data easily drift out of sync between the two environments |
| Fits changes that can't scale up gradually | Poor for validating quality differences (no same-traffic control) |
text
┌──────────────┐ traffic ┌──────────────────┐
user ───▶│ Gateway/LB │────────────▶│ Green (old v1) │
└──────────────┘ └──────────────────┘
│ at the switch
└────────────────────────▶ ┌──────────────────┐
│ Blue (new v2) │
└──────────────────┘When blue/green makes sense: infrastructure changes (major dependency upgrades, CUDA upgrades, K8s cluster migrations) — changes you can't validate with 5% of traffic, only by standing up a whole environment. The price is cost. The decision rule: which is more expensive, the "full-switch risk" or "double the cost"?
5. A/B Testing: Splitting, Significance, and Duration
A canary answers "can we ship it"; A/B answers "is it actually better." The two are often combined: after the canary reaches full rollout, an A/B test validates the algorithm's quality over the long term.
5.1 Traffic Splitting and the Control Group
- Splitting unit: hash on user ID so the same user always lands in the same version (avoiding "v2 today, v1 tomorrow" polluting your results);
- Sample size: pre-compute the minimum sample size (statistical tools help — e.g. Evan Miller's sample size calculator); rule of thumb: lifting CTR from 3% to 3.2% (+6.7%) with a 50/50 split takes roughly 500k impressions for 95% confidence;
- No contamination: don't stack other changes on top during the A/B ("change one variable at a time").
5.2 Significance and Duration
text
Decision criteria (three escalating gates):
1. Confidence interval excludes 0 (the effect is statistically significant)
2. Lift > minimum business value (statistically significant ≠ worth shipping)
3. No unexpected side effects (related metrics haven't worsened, e.g. "CTR up but return rate up too")
What determines duration:
- Enough samples (a low-traffic model may need weeks)
- Covers at least one full business cycle (weekly/monthly rhythm)
- Seasonality ruled out: compare against the baseline over the same periodA/B Is Not a Release Tool
During an A/B test, real users stay on the worse version for the whole duration. If the gap is obvious, you're experimenting on real people. So pair A/B with loss-cutting rules (e.g. if metrics clearly degrade after a 3-day observation window, roll back to 100% immediately) — don't just "wait out the full 4 weeks."
6. Rollback Mechanics: Three Tiers
Rollback isn't just "switch back to the old model." There are three distinct tiers, each with its own rollback path:
| Tier | What rolls back | How | Speed |
|---|---|---|---|
| Model tier | Model version | Point the model registry at the old version (MLflow etc.); the gateway/service reloads | Seconds to minutes |
| Feature tier | Feature/preprocessing logic | Roll the feature service back to the old feature config; model and feature versions must be paired | Minutes |
| Code tier | Serving code | Roll back to the old image / old deployment | Minutes |
The Three Tiers Must Move Together
The model was trained against the "old features." Rolling back the model without rolling back the features is like new shoes with old socks — the results are still wrong. Before rolling back, confirm whether the model, feature, and code versions are still a combination that was trained and validated together. For model version management, see The MLOps Pipeline.
bash
# Typical rollback: point a K8s Deployment back to the previous image
kubectl rollout undo deployment/inference-api
# Or at the model tier: update the config to the old version and reload
kubectl patch configmap model-config --type merge \
-p '{"data":{"model_version":"v1.2.0"}}'7. Pre-Release Checklist
- [ ] Offline evaluation complete; the new model hits its validation-set targets;
- [ ] Quality metrics are defined and computed identically for old and new versions (the same calculation logic);
- [ ] The four-stage canary plan is written down: stages, what to watch at each, dwell time;
- [ ] Automatic rollback conditions are configured at the gateway/orchestration layer and rehearsed at least once;
- [ ] The three-tier rollback plan is ready: model version, feature version, and code image all tagged;
- [ ] Dashboards are ready: side-by-side dual-version views plus quality-metric views;
- [ ] Communication is in place: release channel, on-call roster, owner (a release is a team event);
- [ ] Post-rollback actions are planned: root-cause analysis, bug fix, then a fresh canary.
8. Common Pitfalls
- Cache pollution during the canary: old and new versions share a cache, the new model's inputs hit stale entries, and the quality comparison is skewed. Fix: include the model version in cache keys, or disable volatile caches during the rollout.
- Inconsistent quality-metric definitions: the new version's metric is computed at the model layer, the old one's at the business layer — the "comparison" is meaningless. Fix: funnel all metric computation through one service and split only by version.
- Skipping stages: jumping from 5% straight to 100% leaves too many variables to untangle when something breaks. Fix: stages are discipline, not bureaucracy; skipping requires written sign-off.
- A rollback plan that exists but was never rehearsed: when things actually break, you discover the rollback command lacks permissions or the image tag is wrong. Fix: do a dry-run rollback before the release (shift 1% of traffic, then shift it back).
- Watching only system metrics, not quality: P99 looks perfect while conversion drops 5% — the system "healthily gets worse." Fix: make quality metrics part of the release gate.
Checklist
- [ ] Chose a strategy that matches the change type (canary by default);
- [ ] The four-stage traffic plan and automatic rollback conditions are configured and rehearsed;
- [ ] Three-tier rollback (model / features / code) version pairings confirmed;
- [ ] Quality-metric definitions unified, with an old-version control;
- [ ] For A/B: sample size, significance threshold, and loss-cutting rules all ready;
- [ ] Someone is scheduled to keep watching for 24-48 hours after the release.
Further Reading
- Model Gateway and Canary Releases — the complete walkthrough of gateway routing, splitting, and canary releases
- The MLOps Pipeline — model registry, version management, and pipeline context
- Observability in Practice — dual-version monitoring and alerting during a canary
- Monitoring and Observability — the theory behind SLOs and drift detection
- Common Pitfalls and Anti-Patterns — full-rollout disasters and their root causes
- Load Testing and Capacity Planning — capacity validation before release (can the new version survive peak traffic?)
References
- Argo Rollouts (K8s canary/blue-green): https://argoproj.github.io/rollouts/
- Kubernetes docs (Deployment rollback): https://kubernetes.io/docs/concepts/workloads/controllers/deployment/
- Evan Miller's sample size calculator: https://www.evanmiller.org/ab-testing/sample-size.html
- Google SRE Book (release engineering): https://sre.google/sre-book/release-engineering/