Theme
MLOps and Model Deployment
Concept Definition: Keeping Models Alive in Production
MLOps (Machine Learning Operations) encompasses all the engineering practices needed to take a machine learning system from "a notebook in the lab" to "a stable service running in production" — analogous to DevOps in software engineering, but with two additional versioned objects: data and models.
A harsh industry consensus: most ML projects die after going live, not before training. A model going live is just the beginning: data drifts, features become stale, dependencies upgrade, traffic patterns change — without MLOps support, a model is basically a "zombie service" within three months (still responding, but performance has long since degraded from what was originally validated).
Development: Data → Features → Training → Evaluation → Registry
Production: Serving → Monitoring → Drift Detection → Retraining → Update (closed loop)The Four Pillars of MLOps
1. Experiment Management: Every Experiment Reproducible
Training experiments must be fully logged, otherwise the question "what configuration was the best model last time used?" can't be answered. What to log:
- Code version (Git commit), data version (see Data and Data Engineering);
- Hyperparameters, random seeds, training/validation metric curves;
- Artifacts (weight files, tokenizers, evaluation reports).
Tools: MLflow (the open-source de facto standard), Weights & Biases, Neptune. Minimum viable setup: a named experiment directory with well-defined naming convention + a logging file (experiments/2024-01-01_lr1e-3_seed42/) is already better than logging nothing.
2. Model Registry and Version Management
Model versioning: every candidate model gets a unique version number, a status (staging/production), and metadata (training data, metrics, approver). The go-live process: candidate model → offline evaluation → human approval → canary → full rollout.
Tools: MLflow Model Registry, SageMaker Model Registry. Core value: rollbackability — when the live model breaks, you can revert to the previous version in seconds.
3. Pipeline Orchestration
Automate the training pipeline: data validation → feature computation → training → evaluation → registry, every step re-runnable and auditable. Tools: Airflow, Kubeflow Pipelines, Metaflow, Prefect.
Key practice: dual validation of data and models — check data schemas at pipeline entry (are all columns present, types correct, distributions normal?), and alert rather than blindly training when checks fail.
4. Deployment and Inference Services
How models go live as services (see "Deployment Options" below), plus capacity planning (QPS, latency, GPU resources), auto-scaling, and failover.
Deployment Options
| Option | Characteristics | Use Case |
|---|---|---|
| Online API (online inference) | Real-time, millisecond response | Recommendation, risk control, search, chat |
| Batch processing (offline inference) | Runs once per day/hour, results stored | Daily reports, profile updates, offline scoring |
| Streaming inference | Event-driven, second-level | Real-time risk control, monitoring alerts |
| Embedded / Edge | Model goes into App/device | Mobile, IoT, offline-available |
Common implementations for online inference: wrap the model in FastAPI/Flask → Docker → deploy to K8s/Serverless (KServe, SageMaker, etc.). Key points for model serving:
- Preprocessing and postprocessing must be encapsulated in the service (consistent with training, otherwise features are inconsistent);
- Input validation, timeouts, rate limiting, retries;
- Model loading management (keep models in memory/GPU memory, avoid re-loading for each request).
Performance and Cost Optimization
After a model goes live, inference cost and latency are ongoing bills:
| Technique | Principle | Benefit |
|---|---|---|
| Quantization | FP32→FP16/INT8, weights represented with fewer bits | 50%+ less memory, speedup |
| Distillation | Large model teaches small model | Small model approaches large model performance |
| Pruning | Cut unimportant weights/layers | Smaller, faster model |
| Caching | Cache outputs for identical inputs | Direct hit for frequent repeated queries |
| Batch inference | Merge multiple requests into one forward pass | Greatly improved GPU utilization |
| ONNX/TensorRT | Inference engine optimization | Several-fold speedup |
Engineering order: start with quantization (easiest) → add distillation/pruning if needed → add inference engines if needed. The large model era also has KV Cache, speculative decoding, vLLM, and other inference optimizations — see Large Language Models (LLM).
Monitoring: Models "Expire"
Monitoring is the biggest difference between MLOps and DevOps — code doesn't change on its own, but data and patterns do. Monitor three things:
1. Data Drift
Input feature distributions change: the age distribution of users shifts, new product categories launch. Detection methods:
- PSI (Population Stability Index): compare feature distribution shifts (training vs. recent online), trigger alert if PSI > 0.2;
- KS test: are the two distributions significantly different;
- Per-feature monitoring + overall monitoring (embedding distributions).
2. Concept Drift
The x→y relationship has changed: the same feature combinations follow different rules (the "feature → churn" relationship broke during the pandemic, policy changes alter risk control logic). Detection: compare model predictions against actual results as real labels arrive with delay (e.g., click-through rate, conversion rate).
3. System Health
Inference latency, error rate, QPS, GPU utilization — standard service monitoring suffices.
Actions After an Alert: Retraining
| Trigger Condition | Action |
|---|---|
| Scheduled (daily/weekly) | Periodic retraining (easiest, model keeps pace with data) |
| Drift detection exceeds threshold | On-demand retraining (saves resources, fast response) |
| Online metric degradation | Immediate retraining + roll back previous version |
The core discipline of the retraining closed loop: connect online monitoring metrics with offline evaluation metrics — if offline AUC hasn't changed but online conversion has dropped, first check feature consistency (differences between training/online definitions), which is the most common cause.
MLOps Maturity Ladder
| Stage | Characteristics | Tools |
|---|---|---|
| L0 (Manual) | Notebook training, manual deployment, no monitoring | None |
| L1 (Scripted) | Training scripted + version-controlled, scheduled batch processing | Git, cron |
| L2 (Automated) | Pipeline orchestration + experiment management + model registry | MLflow, Airflow |
| L3 (Platform) | Feature store + auto-retraining + full-stack monitoring | Feast, Kubeflow, SageMaker |
Advice for teams: don't try to deploy the full platform all at once — first ensure these three bottom lines: "experiments are reproducible + models are rollbackable + monitoring has alerts," then upgrade step by step. Tools are secondary; processes are primary.
Build three habits from day one
- Always log experiments: store code/data/hyperparams/metrics for every training run;
- Always version models: register before going live, so you can roll back if something breaks;
- Always monitor after deployment: deploying a model without monitoring = running naked. These three habits cost nothing, yet they determine whether your model survives past three months.
Tradeoffs
- Online vs batch: latency-sensitive (recommendation, risk control) → online; not time-sensitive (profiles, reports) → batch saves 90% of costs;
- Model accuracy vs inference cost: large models have high accuracy but are expensive — combine quantization + distillation + caching, choose a tier based on business budget;
- Automation vs control: automated retraining is efficient but can be misled by "dirty data" (do data validation before retraining); keep manual approval for critical business;
- Self-built vs platform: cloud platforms (SageMaker, Azure ML, etc.) are quick to adopt; self-built (K8s + KServe) is flexible but heavy to maintain — small teams should prefer platforms.
Further Reading
- Data and Data Engineering — the upstream of pipelines
- Model Evaluation and Validation — the closed loop of pre-live and post-live evaluation
- Building an Evaluation System from Scratch — landing evaluation methodology
- How to Choose Frameworks and Tools — MLflow/K8s/cloud platform selection
- Common Pitfalls and Anti-Patterns — training/inference inconsistency failure stories
- Large Language Models (LLM) — large model deployment and inference optimization
References
- MLflow official documentation — experiment management/model registry open-source standard
- Kreuzberger, Kühl, Hirschl. Machine Learning Operations (MLOps): Overview, Definition, and Architecture (2022) — MLOps survey
- Polyzotis et al. Data Lifecycle Challenges in Production Machine Learning (SIGMOD 2018) — data lifecycle in production ML
- Google: MLOps: Continuous delivery and automation pipelines in machine learning — authoritative MLOps maturity documentation
- Breck et al. What's your ML Test Score? (NeurIPS 2016) — the classic ML test checklist