Skip to content

MLOps Deployment Pipelines

At a glance Model CI/CD is not the same as code CI/CD: you also have to manage model versions, data versions, and quality regressions. This article walks through the deployment pipeline from commit to production — model registry, test gates, canary releases, and rollback.

MLOps Deployment Pipelines ​

One-line definition: an MLOps deployment pipeline is the automated path that carries a trained model from the research environment into production — covering model registration, build, offline evaluation, test gates, approval, canary rollout, and rollback, so that models can be reproduced, audited, and rolled back just like code.

Industry insight: the scariest part of shipping a model isn't the launch itself — it's being unable to ship and unable to go back. Many teams still run on "training script + manual copy + gut-feel canary," and when something breaks, retraining from scratch is the only option. Google's research suggests that roughly 90% of machine learning systems never actually reach production — the bottleneck isn't model quality, it's engineering capability. The value of a pipeline isn't "shipping faster"; it's making every release reproducible, rollback-able, and accountable.

1. The Four Pillars of MLOps ​

PillarQuestion it answersKey artifacts
Experiment trackingHow was this model trained?Experiment records, hyperparameters, data versions
Model registryWhich model is the current official version?Model registry, version numbers, status
Pipeline orchestrationHow do data, model, and deployment connect end to end?DAGs, scheduling, dependencies
DeploymentHow does the model go live and roll back safely?Release process, canary strategy, monitoring

This article focuses on the deployment side. For the four pillars in full and how teams implement them, see Optimization in Practice and Release Strategies: Canary and Rollback.

2. Model Registry and Versioning ​

1. The Model Registry ​

Like a code repository, a model needs its own "repository + versions + status":

text
Model registry (e.g., MLflow Model Registry / DVC)
 ├─ Model name: ctr-v3
 ├─ Version: v7 (incremented on each train + register run)
 ├─ Status: None → Staging → Production → Archived
 ├─ Metadata: source experiment id, training data version, metrics (AUC, latency), author
 └─ Artifacts: weight files (hash-addressed, immutable), ONNX/engine, preprocessing code version, dependency manifest

The registry is the precondition for rollback: the Production status always points at one explicit version, and rollback simply means flipping that pointer back to the previous version.

2. Version Immutability ​

Once registered, model artifacts are immutable (content-addressed — any change produces a new version). Every environment then loads the exact same artifact. This is the foundation of reproducibility.

3. The Deployment Pipeline: From Commit to Production ​

text
① Model submission
   └─► ② Build (package the image: weights + runtime + preprocessing code + dependencies)
        └─► ③ Offline evaluation (metrics on train / validation / production replay sets)
             └─► ④ Test gates (see next section)
                  └─► ⑤ Approval (optional, human gate)
                       └─► ⑥ Canary (5% → 20% → 100%)
                            └─► ⑦ Full rollout + monitoring takes over
                                 └─► ⑧ Rollback plan on standby

Every stage must leave a record: who submitted, what the evaluation scores were, who approved, how far the canary got — audit logging is one of the biggest differences between MLOps and ordinary CI/CD.

4. Model Testing vs. Code Testing ​

DimensionCode testingModel testing
What you assertBehavior is correct (function outputs)Metrics hit targets (AUC, latency) + structure is correct
Data dependencyNoneRequires datasets (same data versions)
StabilityDeterministicStochastic (needs repeated evaluation)
ScopeLogicData schema, quality regression, slices, adversarial cases

Recommended model test gates (ordered top to bottom by increasing cost):

  1. Data schema validation: live input features match the training schema (fields, types, missing rates);
  2. Quality regression (offline evaluation): the new model is no worse than the current production model on a fixed evaluation set (e.g., AUC ≥ production + 0.002);
  3. Slice tests: evaluate across segments (demographics, regions, time windows) to catch "overall improves, underserved groups collapse";
  4. Adversarial / robustness tests: does quality hold up when inputs are perturbed (noise, format changes);
  5. Latency / resource regression: model size and inference latency stay within budget.

5. Canary Releases and Rollback ​

1. Canary Strategy ​

text
5% → 20% → 100%  soak at each stage (e.g., 15–60 minutes)
Go/no-go signals: P99 latency, error rate, business metrics (e.g., CTR), no new alerts

The key point: a canary is a controlled experiment running on real traffic, so during rollout you must watch system metrics and quality metrics side by side (see Monitoring and Observability). For the full mechanism and tooling (gateway traffic splitting, shadow mode), see Model Gateways and Canary Releases.

2. Rollback Plan ​

  • Immediate rollback: switch to the previous Production version in the registry (in-process hot swap or restart);
  • Rollback is all-or-nothing: when you roll back a model, the preprocessing code, feature versions, and dependencies must roll back with it — a "new preprocessing + old model" mismatch is worse than not rolling back at all;
  • Rollback triggers: after canary or full rollout, error rate above threshold, or a core business metric down beyond threshold → roll back immediately.

6. A/B Test Design ​

A canary is a platform mechanism; an A/B test is a scientific experiment — the goal is a statistical answer to "is the new model really better?":

Design elementRule of thumb
AssignmentHash user_id so each user always lands in the same bucket
Sample sizePre-compute with a power analysis (small effect size → more samples)
Significancep < 0.05, and check confidence intervals on business metrics
DurationCover a full business cycle (e.g., 7 days including a weekend)
PitfallsWatch for the novelty effect (early numbers look inflated) and cross-contamination

7. Maturity Ladder: L0–L3 ​

LevelCharacteristicsTeam profile
L0Manual training, manual copy, manual releaseFine during experimentation
L1Experiment tracking + model registry + semi-automated deploymentMaintained by an engineer
L2Fully automated pipeline + test gates + canary and rollbackA real production team
L3Automatic retraining (triggered by data drift) + quality feedback loopHigh maturity; see drift detection in Monitoring and Observability

Most teams should stop at L2: L3's automatic retraining brings a quality feedback loop plus cost, and not every scenario justifies it.

8. The Minimum Bar for Teams ​

If you can't afford a full pipeline, hold at least these three lines:

text
① Reproducible: model = code + data version + hyperparameters — the three together reproduce the same weights
② Rollback-able: Production versions are recorded; you can switch back within 10 minutes of an incident
③ Monitored: golden signals + quality metrics live in production, and someone responds to alerts

These three are the baseline for keeping model incidents contained. If any one of them is missing, don't ship.

Trade-offs ​

DecisionOptionsHow to choose
Automation levelManual vs. fully automatedMatch team size and release cadence; semi-automated first, then fully automated
Test depthLightweight vs. heavyHeavy for high-risk models (risk control, healthcare); light for internal tools
Canary step size5% vs. 20%Start at 5% when the quality impact is large and hard to evaluate
Rollback methodIn-process hot swap vs. restartHot swap when multiple versions can coexist; otherwise fast restart
Auto-retrainingYes vs. noGo L3 only when drift is clear and the benefit is quantifiable

In one sentence: the goal of an MLOps pipeline isn't speed — it's making every release reproducible, rollback-able, and accountable. From model registration to canary and rollback, it turns shipping from a gamble into routine operations.

Further Reading ​

References ​