Appearance
MLOps Deployment Pipelines
One-line definition: an MLOps deployment pipeline is the automated path that carries a trained model from the research environment into production — covering model registration, build, offline evaluation, test gates, approval, canary rollout, and rollback, so that models can be reproduced, audited, and rolled back just like code.
Industry insight: the scariest part of shipping a model isn't the launch itself — it's being unable to ship and unable to go back. Many teams still run on "training script + manual copy + gut-feel canary," and when something breaks, retraining from scratch is the only option. Google's research suggests that roughly 90% of machine learning systems never actually reach production — the bottleneck isn't model quality, it's engineering capability. The value of a pipeline isn't "shipping faster"; it's making every release reproducible, rollback-able, and accountable.
1. The Four Pillars of MLOps
| Pillar | Question it answers | Key artifacts |
|---|---|---|
| Experiment tracking | How was this model trained? | Experiment records, hyperparameters, data versions |
| Model registry | Which model is the current official version? | Model registry, version numbers, status |
| Pipeline orchestration | How do data, model, and deployment connect end to end? | DAGs, scheduling, dependencies |
| Deployment | How does the model go live and roll back safely? | Release process, canary strategy, monitoring |
This article focuses on the deployment side. For the four pillars in full and how teams implement them, see Optimization in Practice and Release Strategies: Canary and Rollback.
2. Model Registry and Versioning
1. The Model Registry
Like a code repository, a model needs its own "repository + versions + status":
text
Model registry (e.g., MLflow Model Registry / DVC)
├─ Model name: ctr-v3
├─ Version: v7 (incremented on each train + register run)
├─ Status: None → Staging → Production → Archived
├─ Metadata: source experiment id, training data version, metrics (AUC, latency), author
└─ Artifacts: weight files (hash-addressed, immutable), ONNX/engine, preprocessing code version, dependency manifestThe registry is the precondition for rollback: the Production status always points at one explicit version, and rollback simply means flipping that pointer back to the previous version.
2. Version Immutability
Once registered, model artifacts are immutable (content-addressed — any change produces a new version). Every environment then loads the exact same artifact. This is the foundation of reproducibility.
3. The Deployment Pipeline: From Commit to Production
text
① Model submission
└─► ② Build (package the image: weights + runtime + preprocessing code + dependencies)
└─► ③ Offline evaluation (metrics on train / validation / production replay sets)
└─► ④ Test gates (see next section)
└─► ⑤ Approval (optional, human gate)
└─► ⑥ Canary (5% → 20% → 100%)
└─► ⑦ Full rollout + monitoring takes over
└─► ⑧ Rollback plan on standbyEvery stage must leave a record: who submitted, what the evaluation scores were, who approved, how far the canary got — audit logging is one of the biggest differences between MLOps and ordinary CI/CD.
4. Model Testing vs. Code Testing
| Dimension | Code testing | Model testing |
|---|---|---|
| What you assert | Behavior is correct (function outputs) | Metrics hit targets (AUC, latency) + structure is correct |
| Data dependency | None | Requires datasets (same data versions) |
| Stability | Deterministic | Stochastic (needs repeated evaluation) |
| Scope | Logic | Data schema, quality regression, slices, adversarial cases |
Recommended model test gates (ordered top to bottom by increasing cost):
- Data schema validation: live input features match the training schema (fields, types, missing rates);
- Quality regression (offline evaluation): the new model is no worse than the current production model on a fixed evaluation set (e.g., AUC ≥ production + 0.002);
- Slice tests: evaluate across segments (demographics, regions, time windows) to catch "overall improves, underserved groups collapse";
- Adversarial / robustness tests: does quality hold up when inputs are perturbed (noise, format changes);
- Latency / resource regression: model size and inference latency stay within budget.
5. Canary Releases and Rollback
1. Canary Strategy
text
5% → 20% → 100% soak at each stage (e.g., 15–60 minutes)
Go/no-go signals: P99 latency, error rate, business metrics (e.g., CTR), no new alertsThe key point: a canary is a controlled experiment running on real traffic, so during rollout you must watch system metrics and quality metrics side by side (see Monitoring and Observability). For the full mechanism and tooling (gateway traffic splitting, shadow mode), see Model Gateways and Canary Releases.
2. Rollback Plan
- Immediate rollback: switch to the previous Production version in the registry (in-process hot swap or restart);
- Rollback is all-or-nothing: when you roll back a model, the preprocessing code, feature versions, and dependencies must roll back with it — a "new preprocessing + old model" mismatch is worse than not rolling back at all;
- Rollback triggers: after canary or full rollout, error rate above threshold, or a core business metric down beyond threshold → roll back immediately.
6. A/B Test Design
A canary is a platform mechanism; an A/B test is a scientific experiment — the goal is a statistical answer to "is the new model really better?":
| Design element | Rule of thumb |
|---|---|
| Assignment | Hash user_id so each user always lands in the same bucket |
| Sample size | Pre-compute with a power analysis (small effect size → more samples) |
| Significance | p < 0.05, and check confidence intervals on business metrics |
| Duration | Cover a full business cycle (e.g., 7 days including a weekend) |
| Pitfalls | Watch for the novelty effect (early numbers look inflated) and cross-contamination |
7. Maturity Ladder: L0–L3
| Level | Characteristics | Team profile |
|---|---|---|
| L0 | Manual training, manual copy, manual release | Fine during experimentation |
| L1 | Experiment tracking + model registry + semi-automated deployment | Maintained by an engineer |
| L2 | Fully automated pipeline + test gates + canary and rollback | A real production team |
| L3 | Automatic retraining (triggered by data drift) + quality feedback loop | High maturity; see drift detection in Monitoring and Observability |
Most teams should stop at L2: L3's automatic retraining brings a quality feedback loop plus cost, and not every scenario justifies it.
8. The Minimum Bar for Teams
If you can't afford a full pipeline, hold at least these three lines:
text
① Reproducible: model = code + data version + hyperparameters — the three together reproduce the same weights
② Rollback-able: Production versions are recorded; you can switch back within 10 minutes of an incident
③ Monitored: golden signals + quality metrics live in production, and someone responds to alertsThese three are the baseline for keeping model incidents contained. If any one of them is missing, don't ship.
Trade-offs
| Decision | Options | How to choose |
|---|---|---|
| Automation level | Manual vs. fully automated | Match team size and release cadence; semi-automated first, then fully automated |
| Test depth | Lightweight vs. heavy | Heavy for high-risk models (risk control, healthcare); light for internal tools |
| Canary step size | 5% vs. 20% | Start at 5% when the quality impact is large and hard to evaluate |
| Rollback method | In-process hot swap vs. restart | Hot swap when multiple versions can coexist; otherwise fast restart |
| Auto-retraining | Yes vs. no | Go L3 only when drift is clear and the benefit is quantifiable |
In one sentence: the goal of an MLOps pipeline isn't speed — it's making every release reproducible, rollback-able, and accountable. From model registration to canary and rollback, it turns shipping from a gamble into routine operations.
Further Reading
- Monitoring and Observability — post-launch quality monitoring and drift detection
- Release Strategies: Canary and Rollback — the full mechanics of canary rollout and rollback
- Model Gateways and Canary Releases — hands-on A/B testing with gateway traffic splitting
- Common Pitfalls and Anti-Patterns — classic incidents caused by missing release processes
- Choosing Frameworks and Platforms — selecting your pipeline toolchain