Skip to content

MLOps and Model Deployment

Quick overview The model finishing training is only the beginning — deployment, monitoring, and retraining are where production ML is truly won or lost. This article covers the full MLOps landscape: experiment management, model registry, deployment options, inference services, drift monitoring, and the retraining loop, plus "why so many models die after going live."

MLOps and Model Deployment ​

Concept Definition: Keeping Models Alive in Production ​

MLOps (Machine Learning Operations) encompasses all the engineering practices needed to take a machine learning system from "a notebook in the lab" to "a stable service running in production" — analogous to DevOps in software engineering, but with two additional versioned objects: data and models.

A harsh industry consensus: most ML projects die after going live, not before training. A model going live is just the beginning: data drifts, features become stale, dependencies upgrade, traffic patterns change — without MLOps support, a model is basically a "zombie service" within three months (still responding, but performance has long since degraded from what was originally validated).

Development:  Data → Features → Training → Evaluation → Registry
Production:   Serving → Monitoring → Drift Detection → Retraining → Update (closed loop)

The Four Pillars of MLOps ​

1. Experiment Management: Every Experiment Reproducible ​

Training experiments must be fully logged, otherwise the question "what configuration was the best model last time used?" can't be answered. What to log:

  • Code version (Git commit), data version (see Data and Data Engineering);
  • Hyperparameters, random seeds, training/validation metric curves;
  • Artifacts (weight files, tokenizers, evaluation reports).

Tools: MLflow (the open-source de facto standard), Weights & Biases, Neptune. Minimum viable setup: a named experiment directory with well-defined naming convention + a logging file (experiments/2024-01-01_lr1e-3_seed42/) is already better than logging nothing.

2. Model Registry and Version Management ​

Model versioning: every candidate model gets a unique version number, a status (staging/production), and metadata (training data, metrics, approver). The go-live process: candidate model → offline evaluation → human approval → canary → full rollout.

Tools: MLflow Model Registry, SageMaker Model Registry. Core value: rollbackability — when the live model breaks, you can revert to the previous version in seconds.

3. Pipeline Orchestration ​

Automate the training pipeline: data validation → feature computation → training → evaluation → registry, every step re-runnable and auditable. Tools: Airflow, Kubeflow Pipelines, Metaflow, Prefect.

Key practice: dual validation of data and models — check data schemas at pipeline entry (are all columns present, types correct, distributions normal?), and alert rather than blindly training when checks fail.

4. Deployment and Inference Services ​

How models go live as services (see "Deployment Options" below), plus capacity planning (QPS, latency, GPU resources), auto-scaling, and failover.

Deployment Options ​

OptionCharacteristicsUse Case
Online API (online inference)Real-time, millisecond responseRecommendation, risk control, search, chat
Batch processing (offline inference)Runs once per day/hour, results storedDaily reports, profile updates, offline scoring
Streaming inferenceEvent-driven, second-levelReal-time risk control, monitoring alerts
Embedded / EdgeModel goes into App/deviceMobile, IoT, offline-available

Common implementations for online inference: wrap the model in FastAPI/Flask → Docker → deploy to K8s/Serverless (KServe, SageMaker, etc.). Key points for model serving:

  • Preprocessing and postprocessing must be encapsulated in the service (consistent with training, otherwise features are inconsistent);
  • Input validation, timeouts, rate limiting, retries;
  • Model loading management (keep models in memory/GPU memory, avoid re-loading for each request).

Performance and Cost Optimization ​

After a model goes live, inference cost and latency are ongoing bills:

TechniquePrincipleBenefit
QuantizationFP32→FP16/INT8, weights represented with fewer bits50%+ less memory, speedup
DistillationLarge model teaches small modelSmall model approaches large model performance
PruningCut unimportant weights/layersSmaller, faster model
CachingCache outputs for identical inputsDirect hit for frequent repeated queries
Batch inferenceMerge multiple requests into one forward passGreatly improved GPU utilization
ONNX/TensorRTInference engine optimizationSeveral-fold speedup

Engineering order: start with quantization (easiest) → add distillation/pruning if needed → add inference engines if needed. The large model era also has KV Cache, speculative decoding, vLLM, and other inference optimizations — see Large Language Models (LLM).

Monitoring: Models "Expire" ​

Monitoring is the biggest difference between MLOps and DevOps — code doesn't change on its own, but data and patterns do. Monitor three things:

1. Data Drift ​

Input feature distributions change: the age distribution of users shifts, new product categories launch. Detection methods:

  • PSI (Population Stability Index): compare feature distribution shifts (training vs. recent online), trigger alert if PSI > 0.2;
  • KS test: are the two distributions significantly different;
  • Per-feature monitoring + overall monitoring (embedding distributions).

2. Concept Drift ​

The x→y relationship has changed: the same feature combinations follow different rules (the "feature → churn" relationship broke during the pandemic, policy changes alter risk control logic). Detection: compare model predictions against actual results as real labels arrive with delay (e.g., click-through rate, conversion rate).

3. System Health ​

Inference latency, error rate, QPS, GPU utilization — standard service monitoring suffices.

Actions After an Alert: Retraining ​

Trigger ConditionAction
Scheduled (daily/weekly)Periodic retraining (easiest, model keeps pace with data)
Drift detection exceeds thresholdOn-demand retraining (saves resources, fast response)
Online metric degradationImmediate retraining + roll back previous version

The core discipline of the retraining closed loop: connect online monitoring metrics with offline evaluation metrics — if offline AUC hasn't changed but online conversion has dropped, first check feature consistency (differences between training/online definitions), which is the most common cause.

MLOps Maturity Ladder ​

StageCharacteristicsTools
L0 (Manual)Notebook training, manual deployment, no monitoringNone
L1 (Scripted)Training scripted + version-controlled, scheduled batch processingGit, cron
L2 (Automated)Pipeline orchestration + experiment management + model registryMLflow, Airflow
L3 (Platform)Feature store + auto-retraining + full-stack monitoringFeast, Kubeflow, SageMaker

Advice for teams: don't try to deploy the full platform all at once — first ensure these three bottom lines: "experiments are reproducible + models are rollbackable + monitoring has alerts," then upgrade step by step. Tools are secondary; processes are primary.

Build three habits from day one

  1. Always log experiments: store code/data/hyperparams/metrics for every training run;
  2. Always version models: register before going live, so you can roll back if something breaks;
  3. Always monitor after deployment: deploying a model without monitoring = running naked. These three habits cost nothing, yet they determine whether your model survives past three months.

Tradeoffs ​

  • Online vs batch: latency-sensitive (recommendation, risk control) → online; not time-sensitive (profiles, reports) → batch saves 90% of costs;
  • Model accuracy vs inference cost: large models have high accuracy but are expensive — combine quantization + distillation + caching, choose a tier based on business budget;
  • Automation vs control: automated retraining is efficient but can be misled by "dirty data" (do data validation before retraining); keep manual approval for critical business;
  • Self-built vs platform: cloud platforms (SageMaker, Azure ML, etc.) are quick to adopt; self-built (K8s + KServe) is flexible but heavy to maintain — small teams should prefer platforms.

Further Reading ​

References ​