Theme
MLOps and Model Deployment
One-sentence definition: MLOps (Machine Learning Operations) encompasses all engineering practices that transform a "working training model" into a "stable, reliable, monitorable, and iterable service." The particularity of deep learning models lies in the fact that they are not just code — they are a trinity of "code + data + weights," making lifecycle management far more complex than in traditional software.
1. The DL Lifecycle Panorama
A deep learning project typically goes through these stages (repeated in cycles) from idea to product:
Problem definition → Data engineering → Experiment/training → Evaluation/model selection → Deployment/inference → Monitoring/feedback → RetrainingKey differences from traditional software:
- The model "grows from the data" — when data distribution shifts, the model "expires."
- Performance cannot be statically guaranteed — validation-set scores before launch ≠ production performance; continuous monitoring is essential.
- Performance depends on the environment — GPU drivers, framework versions, and precision modes can all change behavior.
Evaluation methodologies for each stage of the lifecycle are in Deep Learning Evaluation and Experiments, the data side in Data and Data Engineering; this article focuses on the engineering chain "after training."
2. Experiment Tracking and Model Registry
Experiments without records are as good as not done. Two core components:
Experiment tracking: record the full context of every training run — code version (git commit), data version, hyperparameters, random seed, per-epoch metrics, training logs, and generated visualizations. Mainstream tools:
- Weights & Biases (W&B): interactive dashboards, sweep hyperparameter search (commonly used with Bayesian search in Optimization and Gradient Descent), team collaboration friendly.
- MLflow: open-source, with built-in Tracking + Model Registry + Serving; the top choice for local/self-hosted setups.
- Lightweight options:
neptune.ai,TensorBoard+ custom JSON logging.
Model registry: a controlled registration for "models that passed evaluation" — who, when, which commit, what metrics, what artifact path, supporting "one-click rollback to previous version." The model registry is the "trusted anchor" between deployment and rollback (see below).
Experiment discipline
Fixed seeds, changing one variable per experiment, pre-registering primary metrics — all discussed in Deep Learning Evaluation and Experiments; MLOps just turns these disciplines into infrastructure.
3. Deployment Modes: Online Services, Batch Processing, Edge
| Mode | Characteristics | Use cases | Notable frameworks |
|---|---|---|---|
| Online (real-time) | Low latency, high concurrency, request-response | Recommendations, chatbots, OCR, content moderation | TorchServe, Triton, vLLM, FastAPI |
| Batch (offline) | Scheduled/event-driven, throughput-optimized | Offline recommendations, reports, data cleaning | Airflow, Spark, Ray |
| Edge (on-device) | Runs on-device, offline, privacy-preserving, no network | Mobile, cameras, IoT | ONNX Runtime, TFLite, Core ML |
Core selection dimensions: latency requirements, throughput, compute location, data sensitivity. Latency budgets (p50/p99) for online inference determine whether to use continuous batching engines like vLLM; privacy-sensitive data (healthcare, finance) favors edge deployment.
4. Inference Optimization: Quantization, Distillation, Pruning
A trained model that "works" and one that "runs fast and affordably" are two different things. Four major inference optimization methods:
- Quantization: reduce FP32 weights/activations to FP16/BF16/INT8/INT4, trading fewer bytes for faster speed and smaller GPU memory. Two approaches:
- PTQ (Post-Training Quantization): quantize directly using calibration data, no retraining needed — fast but potentially larger accuracy loss.
- QAT (Quantization-Aware Training): simulate quantization errors during training for minimal accuracy loss, but requires retraining.
- Knowledge distillation: use a large model (teacher) to supervise a small model (student) during training — "having the student learn the teacher's judgments" — approximating teacher capability on a smaller model.
- Pruning: remove unimportant weights/channels/heads (sparsification) to reduce computation; structured pruning (entire channels/heads) accelerates speed more directly than unstructured pruning.
- Compiler optimization: ONNX (cross-framework open format) + TensorRT (NVIDIA's specialized inference engine) for operator fusion, graph optimization, and automatic kernel selection. Inference engines like vLLM further optimize with paged KV cache (see Attention) and continuous batching.
Optimization combo: quantize first (INT8) → distill if accuracy doesn't meet requirements → compile (TensorRT/ONNX Runtime). After each step, verify that accuracy loss remains acceptable on the validation set (evaluation criteria in the evaluation chapter).
5. Monitoring and Data/Concept Drift
After deployment, the most insidious risk is "the world changed but the model didn't." Two types of drift:
- Data drift: the distribution of online inputs deviates from the training distribution (new user groups, new languages, seasonal changes). Detection: feature distribution monitoring (KS test, PSI metric), similarity comparison of online vs. training samples.
- Concept drift: the input distribution hasn't changed, but the "input → label" relationship has (e.g., the meaning of "purchase behavior" shifts with the economic environment). Detection: label rate changes, offline backtesting of online samples (shadow deployment).
Monitoring and alerting system: set thresholds + auto-alerts for key metrics (prediction confidence, class distribution, business metrics), and integrate a manual review loop. Actions after drift is triggered: rollback, supplement training data (Data and Data Engineering), trigger retraining pipelines.
Don't just watch model metrics
What you really need to monitor in production is business metrics (conversion rate, retention, satisfaction); model metrics (accuracy, latency) are just intermediate variables. A disconnect between them often signals a mismatch between evaluation criteria and product objectives.
6. Training Platforms and Resource Scheduling
Infrastructure on the training side:
- GPU cluster scheduling: Kubernetes + schedulers (e.g., Volcano, Kueue) manage training tasks; multi-tenant isolation, queuing, preemption.
- Distributed training: data parallelism (DP/DDP), model parallelism, pipeline parallelism (Pipeline), tensor parallelism, ZeRO (memory optimization) — standard for large model training, connected with the batch/memory discussions in Optimization and Gradient Descent. Specific parallel approaches for LLM training are in Large Language Models (LLM).
- Observability: GPU utilization, memory, network bandwidth, and data loading bottlenecks (see the data engineering chapter) should all be visualized; otherwise, "underutilized GPUs" is pure money burned.
7. CI/CD and Rollback
Models also need "continuous integration / continuous deployment":
- CI (Continuous Integration): automatically run on every code/data change: lint → unit tests → small-scale training smoke test (train a few steps without NaN, loss decreases) → evaluation gate (metrics must not fall below baseline). Baseline discipline for evaluation gates is in the evaluation chapter.
- CD (Continuous Deployment): model registry → canary release (new model serves only 5% of traffic) → compare old vs. new metrics → full rollout or rollback.
- Rollback: model registry + versioned inference services (A/B dual-version coexistence) make rollback a one-click operation; rollback strategies (canary, blue-green, shadow) are selected based on business risk.
8. Cost and Latency Trade-offs
The cost structure of deep learning systems: training cost + inference cost, with inference cost often being the larger, ongoing expense.
- Throughput vs. latency: vLLM's continuous batching maximizes throughput, but single-request latency may be slightly higher; real-time interaction scenarios need p99 latency guarantees, while batch scenarios care purely about throughput.
- Precision vs. cost: INT4 quantization can fit a 70B model on a single GPU, but carries precision/calibration risk (larger models lean more on INT4, see Large Language Models (LLM)); distilled small models are cheap to serve but require extra compute for training.
- Self-hosted vs. managed: self-hosted GPU clusters have high upfront costs and difficult utilization management; cloud-managed (pay-per-use) is flexible but more expensive per unit of compute. Decide based on "whether the load is predictable."
- Monitoring cost: logs, metrics, and sampling all cost money — prioritize based on business impact; don't dump everything in at once.
Trade-offs
Model quality vs. inference cost: this is the core balance of MLOps — large models are accurate but expensive, small models are cheap but may miss detections. The standard approach: use the large model first to validate the upper bound, then approximate with distillation/quantization, and defend the lower bound with evaluation gates.
Latency vs. quality: in real-time scenarios, accept slightly worse models to control latency; in offline batch, stack quality. Don't serve all scenarios with the same configuration.
Fast iteration vs. stability: canary releases let you have both "fast" and "stable"; but every release carries risk. Pre-register primary metrics + gates as guardrails (Training Recipes and Hyperparameter Tuning).
MLOps turns deep learning from "scientific experimentation" into "an operational engineering system" — it doesn't directly improve model accuracy, but it determines whether that accuracy can be continuously, reliably, and affordably converted into business value. For a tool landscape and open-source project listings, see Curated Resource Lists; for framework selection (PyTorch/TensorFlow/JAX, etc.), see How to Choose Frameworks and Tools.
Further Reading
- How to Choose Frameworks and Tools — framework and deployment tool matrix
- Curated Resource Lists — MLOps tool open-source listings
- Debugging and Diagnostics — production issues and inference troubleshooting
- Step-by-Step Tutorials — three deployment practice walkthroughs
- Interpretability and Fairness — post-deployment fairness and monitoring audits
- Portfolio Projects — portfolio projects from a deployment perspective