Appearance
Deploying a model from scratch means walking the complete loop — train → export → package → serve → containerize → load test → monitor — via the smallest viable path. The concept pages explain what inference, model formats, and serving are, but actually getting a small model running as an online service with your own hands is a different matter — it welds dozens of individual facts into one piece of muscle memory. This page answers two questions: what is the minimal complete set for deployment, and where do the pitfalls hide at each step? Work through it once and you will come away with a reusable deployment skeleton: an inference class with unit tests, a FastAPI service, a Dockerfile, and a set of load-testing and monitoring configs. You can reuse this skeleton whenever you swap models or engines, and it doubles as a portfolio piece that holds up on a resume (see Resume Analysis and Packaging for the portfolio angle).
Prerequisites
This guide assumes you know basic Python, can run PyTorch locally, and have Docker on your machine. A GPU is not required — the model in the examples is small enough to run on CPU. The whole process takes about 2-3 hours.
1. The Big Picture: A Seven-Step Loop
Deployment is not a one-shot "wrap the model in an endpoint" job — it is seven steps that interlock. Skip any one of them and production will break in its own characteristic way: without export you stay chained to Python; without packaging, pre- and post-processing drift from training; without load testing you don't know your capacity; without monitoring you are flying blind (for a full teardown of these traps see Common Pitfalls and Anti-Patterns).
text
Train Export Package Serve Containerize Load test Monitor
┌─────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐
│ PyTorch │──▶│ ONNX model │──▶│ Inference │──▶│ FastAPI │──▶│ Docker │──▶│ wrk/locust │──▶│ Prometheus │
│ training│ │ (.onnx) │ │ class (pre/ │ │ service │ │ image │ │ load tests │ │ metrics + │
└─────────┘ │ + meta.json │ │ postprocess) │ │ │ │ │ │ │ │ alerts │
└──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘
step1 step2 step3 step4 step5 step6 step7
accuracy reproducible consistency accessibility portability capacity observability
baseline delivery planningEach step answers one question:
| Step | Deliverable | Question it answers | Acceptance criteria |
|---|---|---|---|
| 1. Train | Weight file | What is the model's accuracy baseline? | A clear validation metric exists, recorded as the "offline baseline" |
| 2. Export | ONNX + metadata | Does it run on a different engine? | Imports cleanly; input/output shapes are fixed |
| 3. Package | Inference class | Do pre- and post-processing match training? | Unit tests pass; shares one copy of the preprocessing code with training |
| 4. Serve | REST API | How do others call it? | Health check + error semantics + documentation |
| 5. Containerize | Image + compose | Does it run anywhere? | One command brings the service up |
| 6. Load test | Capacity report | How much QPS can it take? | An SLO exists and a conclusion is documented |
| 7. Monitor | Metrics + alerts | Who notices first when something breaks? | All golden signals covered |
2. The Project for This Guide
We are deploying a CIFAR-10 image classifier: MobileNetV2 fine-tuned for a few epochs on CIFAR-10 (10 classes, 32×32 thumbnails), landing at roughly 80% accuracy. Why this model:
- Small but real: MobileNetV2's ONNX weights are about 14MB (FP32), and single-threaded CPU inference runs 5-15ms — enough to produce meaningful load-test numbers without a GPU;
- Preprocessing has real complexity: normalization, resizing, channel order — exactly what exposes the #1 pitfall, a training/serving preprocessing mismatch;
- Clean inputs and outputs: image bytes in, top-5 labels out — a natural fit for demonstrating an HTTP service.
The repository layout (each step below fills in more of this tree):
text
image-classifier/
├── train_export.py # 1. Train + export ONNX
├── requirements.txt
├── app/
│ ├── main.py # 4. FastAPI service
│ ├── inference.py # 3. Inference class
│ └── metrics.py # 7. Prometheus metrics
├── models/
│ ├── model.onnx # 2. Export artifact
│ └── meta.json # Label table, preprocessing params
├── tests/
│ └── test_inference.py # Unit tests for the inference class
├── Dockerfile # 5. Containerization
├── docker-compose.yml
└── prometheus.yml # 7. Monitoring config3. Step 1: Train and Export to ONNX
Training is not the focus here, but the last thing you do before exporting matters: record the offline baseline metric. It is the anchor for every later "has production regressed?" comparison (see the baseline measurement section of Model Optimization in Practice).
python
# train_export.py — train + export ONNX; runs on CPU in about 15 minutes
import json
import torch
import torch.nn as nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms, models
# Training and inference MUST share the same preprocessing! Define it once here;
# it gets written into meta.json after export
mean, std = (0.4914, 0.4822, 0.4465), (0.2470, 0.2435, 0.2616)
train_tf = transforms.Compose([
transforms.RandomCrop(32, padding=4),
transforms.RandomHorizontalFlip(),
transforms.ToTensor(),
transforms.Normalize(mean, std),
])
eval_tf = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize(mean, std),
])
def main():
torch.manual_seed(42)
ds_train = datasets.CIFAR10("./data", train=True, download=True, transform=train_tf)
ds_eval = datasets.CIFAR10("./data", train=False, download=True, transform=eval_tf)
model = models.mobilenet_v2(weights=None)
model.classifier[1] = nn.Linear(model.last_channel, 10) # 10 classes
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
for epoch in range(3):
model.train()
for x, y in DataLoader(ds_train, batch_size=128, shuffle=True, num_workers=2):
opt.zero_grad()
nn.functional.cross_entropy(model(x), y).backward()
opt.step()
model.eval()
correct = total = 0
with torch.no_grad():
for x, y in DataLoader(ds_eval, batch_size=256, num_workers=2):
correct += (model(x).argmax(1) == y).sum().item()
total += y.size(0)
acc = correct / total
print(f"Offline baseline accuracy = {acc:.4f}") # expect ~0.79-0.82
# ---------- Export to ONNX ----------
dummy = torch.randn(1, 3, 32, 32)
torch.onnx.export(
model, dummy, "models/model.onnx",
input_names=["input"], output_names=["logits"],
dynamic_axes={"input": {0: "batch"}, "logits": {0: "batch"}},
opset_version=17,
)
# Persist the label table and preprocessing params alongside the model, for serving
with open("models/meta.json", "w", encoding="utf-8") as f:
json.dump({"labels": ds_train.classes, "mean": mean, "std": std,
"input_size": [3, 32, 32], "baseline_acc": acc}, f, indent=2)
print("Exported models/model.onnx and models/meta.json")
if __name__ == "__main__":
main()Four key export decisions:
- Use opset_version 17 or higher (the PyTorch 2.x default is fine) — lower versions are missing operator support;
- Turn on a dynamic batch axis with dynamic_axes, otherwise the server can only process one image at a time. You can also tune thread counts later through ONNX Runtime's
session_options.intra_op_num_threads; - mean/std must be written into meta.json — serving-side preprocessing uses the exact same numbers as training. This is the first line of defense against "great offline, broken online";
- Immediately after exporting, run it once through ONNX Runtime and compare the outputs. Numeric differences on the order of 1e-4 are acceptable; a larger gap means the export has an operator problem (see Model Formats and Conversion).
Run torch.onnx.export with check=True before you call it done
Export runs a built-in verification by default: execute the model once in PyTorch, once with the exported graph, and compare the results. Never disable it by accident (do_constant_folding=False is only for when you are certain you need to keep dynamic shapes). If the numbers disagree, the scene of the crime is the export, not production.
4. Steps 2-3: Wrap the Inference Class — Pre/Post-Processing Identical to Training
The inference class is the heart of this skeleton: the service code depends only on the class and never touches the ONNX session directly. Swapping engines (Triton, vLLM) or moving to a quantized build later means editing exactly one file.
python
# app/inference.py — inference class; shares preprocessing params with training
import json
import threading
import numpy as np
import onnxruntime as ort
from PIL import Image
class ImageClassifier:
def __init__(self, onnx_path: str, meta_path: str):
with open(meta_path, encoding="utf-8") as f:
meta = json.load(f)
self.labels = meta["labels"]
self.mean = np.array(meta["mean"], dtype=np.float32).reshape(3, 1, 1)
self.std = np.array(meta["std"], dtype=np.float32).reshape(3, 1, 1)
# One session, created once for the whole process (see pitfall #5: never build a session per request)
self.sess = ort.InferenceSession(onnx_path,
providers=["CPUExecutionProvider"])
# The onnxruntime session is thread-safe and can be shared; the lock keeps
# queueing behavior consistent across threads
self._lock = threading.Lock()
def preprocess(self, image_bytes: bytes) -> np.ndarray:
# Mirrors the training eval_tf step by step: ToTensor + Normalize
img = Image.open(BytesIO(image_bytes)).convert("RGB").resize((32, 32))
arr = np.asarray(img, dtype=np.float32) / 255.0 # ToTensor
arr = arr.transpose(2, 0, 1) # HWC -> CHW
arr = (arr - self.mean) / self.std # Normalize
return arr[None, ...] # (1,3,32,32)
def predict(self, image_bytes: bytes) -> dict:
x = self.preprocess(image_bytes)
with self._lock:
logits = self.sess.run(["logits"], {"input": x})[0]
probs = np.exp(logits - logits.max(axis=1, keepdims=True))
probs = probs / probs.sum(axis=1, keepdims=True) # softmax
top = np.argsort(probs[0])[::-1][:5]
return {"top5": [{"label": self.labels[i], "prob": float(probs[0][i])}
for i in top]}Why preprocessing lives on the server
Training-time augmentation (RandomCrop and friends) belongs only in the training data pipeline, but ToTensor + Normalize is part of the input distribution and must be reproduced server-side. Leaving preprocessing to the client means "the distribution the model sees no longer matches training" — the #1 pitfall in Common Pitfalls and Anti-Patterns.
The matching unit test (3 lines as a placeholder, about 20 in real life):
python
# tests/test_inference.py
def test_preprocess_matches_training():
# Use an image from the training set: hand-compute ToTensor+Normalize and
# compare against preprocess(). Assert np.allclose(...) — this is the
# automated acceptance test for pre/post-processing consistency
...5. Step 4: Serving with FastAPI
The service layer does exactly three things: accept requests, call the inference class, return a structured response. Business logic (retries, authentication, rate limiting) does not belong here — that is the gateway's job (covered in Model Gateway and Canary Releases). For a deeper FastAPI walkthrough see An Online Service with FastAPI + Docker.
python
# app/main.py
import time
from fastapi import FastAPI, File, UploadFile, HTTPException
from app.inference import ImageClassifier
from app.metrics import metrics
app = FastAPI(title="image-classifier")
clf = ImageClassifier("models/model.onnx", "models/meta.json")
@app.get("/healthz")
def healthz():
return {"status": "ok"} # For K8s/Docker health checks; no inference logic here
@app.post("/predict")
def predict(file: UploadFile = File(...)):
data = file.file.read()
t0 = time.perf_counter()
try:
result = clf.predict(data)
except Exception as e: # Corrupt image, etc. — return 400, not 500
raise HTTPException(status_code=400, detail=f"bad image: {e}")
latency_ms = (time.perf_counter() - t0) * 1000
metrics.observe(latency_ms) # 7. Instrumentation: counters, histograms
return {"result": result, "latency_ms": round(latency_ms, 2)}Launch commands (never run a single uvicorn worker in production — see pitfall 5):
bash
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 8000 --workers 4
curl -F "file=@cat.jpg" http://localhost:8000/predict
curl http://localhost:8000/healthzThe single-worker trap
uvicorn --workers 1 is fine for local debugging, but in production one process means one copy of the model and single-threaded request handling. With --workers 4, every worker loads its own copy of the model (4x the memory). This is the memory/VRAM accounting mistake that trips up deployment newcomers most often (Memory Leaks and OOM has the full analysis).
6. Step 5: Dockerize
The goal of containerization is reproduction on any machine with one command. Build the image in two stages — install dependencies first, then copy the model and code — so the model layer and the dependency layer cache separately.
dockerfile
# Dockerfile — built on the slim image, at least half the size
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app/ app/
COPY models/ models/
EXPOSE 8000
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]yaml
# docker-compose.yml — service + monitoring, one command
services:
api:
build: .
ports: ["8000:8000"]
restart: unless-stopped
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/healthz"]
interval: 30s
timeout: 5s
retries: 3
prometheus:
image: prom/prometheus:v2.53.0
volumes: ["./prometheus.yml:/etc/prometheus/prometheus.yml"]
ports: ["9090:9090"]Build and verify:
bash
docker compose up -d --build
docker compose ps # status should read healthy
curl -F "file=@cat.jpg" http://localhost:8000/predictWrite requirements.txt with pinned versions, never >=:
text
fastapi==0.115.6
uvicorn[standard]==0.32.1
onnxruntime==1.19.2
Pillow==11.0.0
numpy==1.26.4
prometheus-client==0.21.17. Step 6: Local Load Testing — Baseline First, Optimization Later
Use wrk to grab a rough baseline (5 seconds, 4 threads, 100 connections). POSTing images is awkward with wrk, and adding a GET /predict_path?url= endpoint just for testing is a bad idea — the simpler route: load test /healthz to get the bare service-layer throughput, then benchmark the inference function separately. The more complete approach is to load test a path with a test image embedded; for tools and methodology see Load Testing and Capacity Planning.
bash
# Bare service-layer throughput (no inference)
wrk -t4 -c100 -d10s http://localhost:8000/healthz
# Read Requests/sec and Latency P99 from the output; record as the "service-layer baseline"
# Real inference path: script it with locust and POST the image file
pip install locustpython
# locustfile.py — load test against the real inference path
import io
from PIL import Image
from locust import HttpUser, task, between
class ImageUser(HttpUser):
wait_time = between(0.05, 0.15) # simulate a 50-150ms gap between requests
img_bytes = io.BytesIO() # one fixed test image at module level, so load-tester IO never becomes the bottleneck
Image.new("RGB", (32, 32), (128, 64, 200)).save(img_bytes, format="PNG")
@task
def predict(self):
self.client.post("/predict", files={"file": ("t.png", self.img_bytes.getvalue())})Three disciplines of load testing (details in Load Testing and Capacity Planning):
- The load generator must not be the bottleneck — keep the test image in memory; never read from disk inside the loop;
- Warm up for at least 30-60 seconds before recording numbers, skipping the inflated readings from JIT/cache warm-up;
- Record the environment: load-generator specs, worker count, concurrency, duration — without these the numbers cannot be reproduced.
8. Step 7: Wiring Up Prometheus Metrics
Use prometheus-client to expose the three golden signals: request counts, a latency histogram, and error counts.
python
# app/metrics.py — the single, global metrics module
from prometheus_client import Counter, Histogram
REQUESTS = Counter("predict_requests_total", "Total inference requests",
["model", "version"])
ERRORS = Counter("predict_errors_total", "Failed inference requests",
["model", "version", "error_type"])
LATENCY = Histogram("predict_latency_seconds",
"Inference latency (seconds)",
buckets=(0.005, 0.01, 0.02, 0.05, 0.1, 0.25, 0.5, 1.0))
def observe(latency_ms: float):
REQUESTS.labels(model="mobilenetv2", version="v1").inc()
LATENCY.observe(latency_ms / 1000.0)Expose /metrics in main.py for Prometheus to scrape:
python
from prometheus_client import generate_latest, CONTENT_TYPE_LATEST
from fastapi import Response
@app.get("/metrics")
def metrics_endpoint():
return Response(generate_latest(), media_type=CONTENT_TYPE_LATEST)yaml
# prometheus.yml
scrape_configs:
- job_name: image-classifier
static_configs:
- targets: ["api:8000"] # use the service name inside the compose network
scrape_interval: 10sOnce it is up (docker compose up -d prometheus), open http://localhost:9090 and query:
text
# Requests per second (5-minute average)
rate(predict_requests_total[5m])
# P99 latency
histogram_quantile(0.99, sum by (le) (rate(predict_latency_seconds_bucket[5m])))At this point the minimal deployment loop is complete: you can train, export, serve, containerize, load test, and observe. For full metric design, alerting rules, and Grafana dashboards, see Putting Observability into Practice.
9. The Go-Live Checklist: Do Not Call It "Deployed" Yet
Tick every item before going to production; if any answer is "No", hold off:
- [ ] The
/healthzendpoint is ready and the container healthcheck is configured; - [ ] Structured logging is wired up (JSON, with a request_id — see Putting Observability into Practice);
- [ ]
/metricscarries the three golden signals and Prometheus scrapes them correctly; - [ ] A rollback plan exists: the old image is retained and one command switches back (see Release Strategies: Canary Releases and Rollbacks);
- [ ] The load test report has conclusions: peak QPS, P99, saturation point;
- [ ] The model-to-code version mapping is recorded (model registry — see The MLOps Pipeline);
- [ ] Memory/VRAM ceilings are measured and there is an OOM contingency plan (Common Pitfalls and Anti-Patterns, pitfalls 4 and 5).
10. Common Pitfalls (the Three Most Likely in This Workflow)
- Preprocessing mismatch: training uses
ToTensor+Normalize, the server forgets the normalization → accuracy drops 20 points and nobody can figure out why. Fix: define preprocessing exactly once intrain_export.py, write it intometa.json, and have the inference class read it from there. - No validation after export: shipping while ONNX and PyTorch outputs disagree. Fix: append an automatic comparison at the end of the export script,
np.allclose(atol=1e-4). - Irreproducible load-test numbers: no environment recorded, no warm-up, load-generator IO as the bottleneck. Fix: put the load-test commands and results into the README (Resume Analysis and Packaging has a template).
11. Where to Go Next: From Skeleton to Production
This skeleton is the foundation; four paths lead upward. Pick based on your goal:
| Direction | How | References |
|---|---|---|
| Performance | INT8 quantization (start with PTQ), switch to a TensorRT engine | Model Optimization in Practice, TensorRT Edge Deployment |
| Multi-model / high throughput | Move to Triton, enable dynamic batching | Multi-Model Serving with NVIDIA Triton |
| LLMs | Move to vLLM: PagedAttention + continuous batching | LLM Serving with vLLM, LLM Inference |
| Scale | K8s + KServe/BentoML, autoscaling | Choosing Between Frameworks and Platforms |
Checklist
- [ ] All seven steps of the loop completed, each with a deliverable and verification;
- [ ] Offline baseline accuracy recorded in meta.json so production metrics can be compared against it;
- [ ] Inference class unit tests pass; preprocessing is identical to training;
- [ ] The container works with a single
docker compose upcommand; - [ ] The load test report includes environment, parameters, and QPS/P99 conclusions;
- [ ] Golden signals are visible in Grafana and every alert has an owner.
Further Reading
- Common Pitfalls and Anti-Patterns — what breaks at each step of this workflow, and why
- Load Testing and Capacity Planning — from a rough baseline to the capacity formula, end to end
- Putting Observability into Practice — complete configuration for metrics, alerting, and logs
- Learning Paths: Three Routes — where this skeleton sits in week 8 of the systematic track
- An Online Service with FastAPI + Docker — a deeper case study on serving
- Resume Analysis and Packaging — how to write this project into your resume
References
- PyTorch ONNX export docs: https://pytorch.org/docs/stable/onnx.html
- ONNX Runtime official docs: https://onnxruntime.ai/docs/
- FastAPI official docs: https://fastapi.tiangolo.com/
- prometheus/client_python: https://github.com/prometheus/client_python
- wrk: https://github.com/wg/wrk
- Locust official docs: https://docs.locust.io/