Appearance
Deployment vs. MLOps vs. Inference: Five Buzzwords, Clearly Separated
Five terms come up daily in the model deployment community and get used interchangeably almost all the time: model deployment, MLOps, inference, model serving, and the inference engine. Ask an interviewee to explain the difference between MLOps and model serving, and fewer than three in ten can sketch the concept hierarchy on the spot. This page nails down the boundaries with a one-line definition for each, a comparison table, and a hierarchy diagram.
The one-sentence version: MLOps ⊃ model deployment ⊃ model serving; inference is a runtime behavior, and the inference engine is the low-level software that executes it.
1. One-Line Definitions of the Five Concepts
| Concept | One-line definition | Boundary/contrast term |
|---|---|---|
| MLOps | The methodology and practice of engineering and automating the entire ML lifecycle (data, training, deployment, monitoring, governance) | DevOps, but for machine learning |
| Model deployment | The engineering process of keeping a trained model effective in a real environment (see What Is Model Deployment?) | Model training |
| Model serving | The technology layer that wraps inference behind a network interface so other systems can call it | Batch/offline inference |
| Inference | A single runtime act of computing a prediction from an input using model parameters | Training (forward vs. backward pass) |
| Inference engine | The low-level software library that executes inference efficiently, such as ONNX Runtime, TensorRT, or vLLM | Naive PyTorch inference via a programming framework |
2. Comparison Table: Definition, Timing, Ownership, and Tools
The same five terms compared along four dimensions:
| Concept | Essence | When it happens | Who owns it | Typical tools/platforms |
|---|---|---|---|---|
| MLOps | Methodology/platform | Entire lifecycle, continuous | MLOps/platform teams | MLflow, Kubeflow, SageMaker, MLOps pipelines |
| Model deployment | Engineering process | After training completes, ongoing | Deployment/inference platform engineers | KServe, Seldon Core, Triton deployments, custom Docker |
| Model serving | Technology layer | The steady state after deployment | Backend/service engineers | FastAPI, Triton, vLLM, KServe |
| Inference | Runtime behavior | On every request | Engine/kernel developers | Forward pass computation |
| Inference engine | Low-level software | On every request | Framework developers/vendors | ONNX Runtime, TensorRT, vLLM, OpenVINO |
A quick self-test: "who owns it?"
The same tool can span multiple concepts, so don't infer the concept from the tool. Triton is a model serving framework that is also commonly used to perform the act of "deployment," but it belongs to the serving layer — it is not an MLOps orchestrator. To classify a tool, look at which problem it solves in the pipeline, not at what its README says.
3. The Hierarchy: What Contains What
These five concepts aren't peers — they nest inside each other:
text
┌──────────────────────────────────────────────────────────┐
│ MLOps (lifecycle-wide methodology) │
│ ┌────────────────────────────────────────────────────┐ │
│ │ Model deployment (keeping models effective │ │
│ │ in the real world) │ │
│ │ ┌──────────────────────────────────────────────┐ │ │
│ │ │ Model serving (inference behind a network │ │ │
│ │ │ interface) │ │ │
│ │ │ ├── FastAPI / gRPC gateway │ │ │
│ │ │ └── Inference engine (executes inference) │ │ │
│ │ │ ├── ONNX Runtime / TensorRT │ │ │
│ │ │ └── vLLM (engine + serving in one, │ │ │
│ │ │ for the LLM era) │ │ │
│ │ └──────────────────────────────────────────────┘ │ │
│ │ ↑ Inference isn't a "layer" — it's a runtime act │ │
│ └────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────┘How to read the diagram:
- MLOps is the largest circle: data management, training, deployment, monitoring, and governance all fall within its scope.
- Deployment is a subset of MLOps: it covers only the "ship and maintain the model" stretch — not data labeling or model design.
- Model serving is a subset of deployment: deployment also handles images, rolling releases, canaries, and rollbacks; serving only cares about "turning requests into predictions."
- Inference isn't a layer — it's an act: it happens inside the serving layer, as one computation whenever the engine is invoked.
- The inference engine sits at the bottom: serving frameworks call engines; engines talk directly to the GPU/CPU.
The most compact way to remember it
MLOps owns the big picture, deployment owns shipping, serving owns the interface, the engine owns the computation — and inference is what happens in that moment.
4. The Full Pipeline: From Trained Model to Production
Placing the five concepts back into a real pipeline makes it obvious where each one shows up:
text
Data prep → Training → Evaluation → Model registry → Image/containerization → Deploy → Serve → Monitor & iterate
│ │ │ │ │ │ │ │
└─────────┴──────────┴────────────┴──────────────────┴───────────────┴─────────┴───────────┘
└────────────── MLOps spans the whole flow ──────────────┘
(MLflow logs experiments and registers models)
Everything after training is "model deployment": │
Dockerize → Orchestrate (K8s/KServe) → Go live → Canary → Monitor │
The "serve" step is model serving: │
A serving framework (FastAPI/Triton) receives an HTTP/gRPC request →│
calls the inference engine (ONNX Runtime/TensorRT) to run one │
"inference" → returns the prediction │Ownership of each step:
| Pipeline step | Owning concept | Typical actions |
|---|---|---|
| Experiment tracking, model version registration | MLOps | MLflow logs metrics, model registry |
| Training, evaluation | MLOps territory (not deployment) | PyTorch / TensorFlow training scripts |
| Exporting the model format | Deployment preparation | Convert to ONNX, model formats |
| Building the serving image, writing the API | Model serving | FastAPI, gRPC, input/output schemas |
| Calling the engine to compute | Inference | ONNX Runtime session run() |
| Container orchestration, canaries, scaling | Model deployment | Kubernetes, KServe deployments |
| Alerting, drift detection, retraining triggers | MLOps | Prometheus + monitoring |
5. Correcting Common Misstatements
These are the mix-ups that show up most often in interviews and community discussions — corrected one by one:
"Deployment means putting the model up once, after training." ✗ Wrong. Deployment is a continuous endeavor: models get updated, environments change, traffic changes. Canary releases, monitoring, rollbacks, and the retraining loop are all part of the full picture (see What Is Model Deployment?).
"An inference engine is the same thing as a serving framework." ✗ Wrong. ONNX Runtime / TensorRT / vLLM execute the computation; FastAPI / Triton / KServe organize requests and resources. Triton blurs the line because it has built-in scheduling for multiple engines, which makes people conflate the two — but its role is still a serving framework.
"Model serving = deployment." ✗ Wrong. Serving is only the "interface layer" of deployment. Deployment also includes image building, rolling releases, rollbacks, and capacity planning — none of which the concept of model serving covers.
"MLOps is just CI/CD plus deployment." ✗ Wrong. MLOps covers data versioning, experiment management, training orchestration, feature engineering, and model governance and compliance; deployment is just one stage in it. See the MLOps pipeline overview.
"Inference = running model(x) in PyTorch." ✗ Half right. Semantically, "computing a prediction with a model" is correct — but in a production context inference means executing efficiently on an engine (batching, quantization, memory reuse). The naive
model(x)call is neither efficient nor production-grade (see Inference fundamentals).
A one-question litmus test
When someone mixes these terms up, ask one counter-question: "Where in the pipeline does that sit, and what problem does it solve?" If they can name the segment and the problem, the concept holds.
6. Why These Five Terms Need to Be Kept Apart
Separating the concepts isn't linguistic pedantry — it has engineering consequences:
- Team boundaries: fuzzy concepts lead to "service engineers writing engine scheduling" or "the platform team not owning API contracts." Ambiguous ownership is a breeding ground for production incidents.
- Tooling decisions: picking tools at the wrong layer means using an MLOps orchestrator to solve a serving-latency problem. See choosing frameworks and platforms.
- Troubleshooting: when a request times out, not knowing whether the cause is "queueing at the serving layer" (a serving problem), "slow engine execution" (an engine problem), or "a scheduling policy issue" (a deployment-layer problem) sends you in circles. Layer-by-layer debugging is covered in Anatomy of an Inference System.
Further Reading
- What Is Model Deployment? — the full definition and challenges of the deployment stage
- MLOps Deployment Pipelines — the full expansion of the "largest circle"
- Model Serving and Inference APIs — serving-layer technology and batching mechanisms
- Inference: From Forward Pass to Inference Engines — the fundamentals of the inference act
- Choosing Frameworks and Platforms — a selection comparison across engines
- Glossary — authoritative definitions for more confusable terms
References
- Google Cloud – MLOps: Continuous delivery and automation pipelines in machine learning
- MLOps.org – The community definition and practice of MLOps
- NVIDIA Triton Inference Server documentation (a serving framework example)
- ONNX Runtime documentation (an inference engine example)
- Machine Learning Systems Design (Chip Huyen's engineering perspective)