Skip to content

Deploy a Model from Scratch

At a glance Walk the complete deployment loop — train → export → serve → load test → monitor — along a minimal viable path. Every step ships with runnable code, a repo layout, common pitfalls, and where to go next, leaving you with a reusable deployment skeleton.

Deploying a model from scratch means walking the complete loop — train → export → package → serve → containerize → load test → monitor — via the smallest viable path. The concept pages explain what inference, model formats, and serving are, but actually getting a small model running as an online service with your own hands is a different matter — it welds dozens of individual facts into one piece of muscle memory. This page answers two questions: what is the minimal complete set for deployment, and where do the pitfalls hide at each step? Work through it once and you will come away with a reusable deployment skeleton: an inference class with unit tests, a FastAPI service, a Dockerfile, and a set of load-testing and monitoring configs. You can reuse this skeleton whenever you swap models or engines, and it doubles as a portfolio piece that holds up on a resume (see Resume Analysis and Packaging for the portfolio angle).

Prerequisites

This guide assumes you know basic Python, can run PyTorch locally, and have Docker on your machine. A GPU is not required — the model in the examples is small enough to run on CPU. The whole process takes about 2-3 hours.

1. The Big Picture: A Seven-Step Loop ​

Deployment is not a one-shot "wrap the model in an endpoint" job — it is seven steps that interlock. Skip any one of them and production will break in its own characteristic way: without export you stay chained to Python; without packaging, pre- and post-processing drift from training; without load testing you don't know your capacity; without monitoring you are flying blind (for a full teardown of these traps see Common Pitfalls and Anti-Patterns).

text
  Train            Export            Package           Serve             Containerize      Load test         Monitor
 ┌─────────┐   ┌──────────────┐   ┌──────────────┐   ┌──────────────┐   ┌──────────────┐   ┌──────────────┐   ┌──────────────┐
 │ PyTorch │──▶│ ONNX model   │──▶│ Inference    │──▶│ FastAPI      │──▶│ Docker       │──▶│ wrk/locust   │──▶│ Prometheus   │
 │ training│   │ (.onnx)      │   │ class (pre/  │   │ service      │   │ image        │   │ load tests   │   │ metrics +    │
 └─────────┘   │ + meta.json  │   │ postprocess) │   │              │   │              │   │              │   │ alerts       │
               └──────────────┘   └──────────────┘   └──────────────┘   └──────────────┘   └──────────────┘   └──────────────┘
   step1          step2              step3              step4              step5             step6              step7
   accuracy       reproducible       consistency        accessibility      portability       capacity           observability
   baseline       delivery                                                                     planning

Each step answers one question:

StepDeliverableQuestion it answersAcceptance criteria
1. TrainWeight fileWhat is the model's accuracy baseline?A clear validation metric exists, recorded as the "offline baseline"
2. ExportONNX + metadataDoes it run on a different engine?Imports cleanly; input/output shapes are fixed
3. PackageInference classDo pre- and post-processing match training?Unit tests pass; shares one copy of the preprocessing code with training
4. ServeREST APIHow do others call it?Health check + error semantics + documentation
5. ContainerizeImage + composeDoes it run anywhere?One command brings the service up
6. Load testCapacity reportHow much QPS can it take?An SLO exists and a conclusion is documented
7. MonitorMetrics + alertsWho notices first when something breaks?All golden signals covered

2. The Project for This Guide ​

We are deploying a CIFAR-10 image classifier: MobileNetV2 fine-tuned for a few epochs on CIFAR-10 (10 classes, 32×32 thumbnails), landing at roughly 80% accuracy. Why this model:

  • Small but real: MobileNetV2's ONNX weights are about 14MB (FP32), and single-threaded CPU inference runs 5-15ms — enough to produce meaningful load-test numbers without a GPU;
  • Preprocessing has real complexity: normalization, resizing, channel order — exactly what exposes the #1 pitfall, a training/serving preprocessing mismatch;
  • Clean inputs and outputs: image bytes in, top-5 labels out — a natural fit for demonstrating an HTTP service.

The repository layout (each step below fills in more of this tree):

text
image-classifier/
├── train_export.py        # 1. Train + export ONNX
├── requirements.txt
├── app/
│   ├── main.py            # 4. FastAPI service
│   ├── inference.py       # 3. Inference class
│   └── metrics.py         # 7. Prometheus metrics
├── models/
│   ├── model.onnx         # 2. Export artifact
│   └── meta.json          # Label table, preprocessing params
├── tests/
│   └── test_inference.py  # Unit tests for the inference class
├── Dockerfile             # 5. Containerization
├── docker-compose.yml
└── prometheus.yml         # 7. Monitoring config

3. Step 1: Train and Export to ONNX ​

Training is not the focus here, but the last thing you do before exporting matters: record the offline baseline metric. It is the anchor for every later "has production regressed?" comparison (see the baseline measurement section of Model Optimization in Practice).

python
# train_export.py — train + export ONNX; runs on CPU in about 15 minutes
import json
import torch
import torch.nn as nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms, models

# Training and inference MUST share the same preprocessing! Define it once here;
# it gets written into meta.json after export
mean, std = (0.4914, 0.4822, 0.4465), (0.2470, 0.2435, 0.2616)
train_tf = transforms.Compose([
    transforms.RandomCrop(32, padding=4),
    transforms.RandomHorizontalFlip(),
    transforms.ToTensor(),
    transforms.Normalize(mean, std),
])
eval_tf = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize(mean, std),
])

def main():
    torch.manual_seed(42)
    ds_train = datasets.CIFAR10("./data", train=True, download=True, transform=train_tf)
    ds_eval  = datasets.CIFAR10("./data", train=False, download=True, transform=eval_tf)
    model = models.mobilenet_v2(weights=None)
    model.classifier[1] = nn.Linear(model.last_channel, 10)  # 10 classes
    opt = torch.optim.Adam(model.parameters(), lr=1e-3)

    for epoch in range(3):
        model.train()
        for x, y in DataLoader(ds_train, batch_size=128, shuffle=True, num_workers=2):
            opt.zero_grad()
            nn.functional.cross_entropy(model(x), y).backward()
            opt.step()

    model.eval()
    correct = total = 0
    with torch.no_grad():
        for x, y in DataLoader(ds_eval, batch_size=256, num_workers=2):
            correct += (model(x).argmax(1) == y).sum().item()
            total += y.size(0)
    acc = correct / total
    print(f"Offline baseline accuracy = {acc:.4f}")  # expect ~0.79-0.82

    # ---------- Export to ONNX ----------
    dummy = torch.randn(1, 3, 32, 32)
    torch.onnx.export(
        model, dummy, "models/model.onnx",
        input_names=["input"], output_names=["logits"],
        dynamic_axes={"input": {0: "batch"}, "logits": {0: "batch"}},
        opset_version=17,
    )
    # Persist the label table and preprocessing params alongside the model, for serving
    with open("models/meta.json", "w", encoding="utf-8") as f:
        json.dump({"labels": ds_train.classes, "mean": mean, "std": std,
                   "input_size": [3, 32, 32], "baseline_acc": acc}, f, indent=2)
    print("Exported models/model.onnx and models/meta.json")

if __name__ == "__main__":
    main()

Four key export decisions:

  1. Use opset_version 17 or higher (the PyTorch 2.x default is fine) — lower versions are missing operator support;
  2. Turn on a dynamic batch axis with dynamic_axes, otherwise the server can only process one image at a time. You can also tune thread counts later through ONNX Runtime's session_options.intra_op_num_threads;
  3. mean/std must be written into meta.json — serving-side preprocessing uses the exact same numbers as training. This is the first line of defense against "great offline, broken online";
  4. Immediately after exporting, run it once through ONNX Runtime and compare the outputs. Numeric differences on the order of 1e-4 are acceptable; a larger gap means the export has an operator problem (see Model Formats and Conversion).

Run torch.onnx.export with check=True before you call it done

Export runs a built-in verification by default: execute the model once in PyTorch, once with the exported graph, and compare the results. Never disable it by accident (do_constant_folding=False is only for when you are certain you need to keep dynamic shapes). If the numbers disagree, the scene of the crime is the export, not production.

4. Steps 2-3: Wrap the Inference Class — Pre/Post-Processing Identical to Training ​

The inference class is the heart of this skeleton: the service code depends only on the class and never touches the ONNX session directly. Swapping engines (Triton, vLLM) or moving to a quantized build later means editing exactly one file.

python
# app/inference.py — inference class; shares preprocessing params with training
import json
import threading
import numpy as np
import onnxruntime as ort
from PIL import Image


class ImageClassifier:
    def __init__(self, onnx_path: str, meta_path: str):
        with open(meta_path, encoding="utf-8") as f:
            meta = json.load(f)
        self.labels = meta["labels"]
        self.mean = np.array(meta["mean"], dtype=np.float32).reshape(3, 1, 1)
        self.std = np.array(meta["std"], dtype=np.float32).reshape(3, 1, 1)
        # One session, created once for the whole process (see pitfall #5: never build a session per request)
        self.sess = ort.InferenceSession(onnx_path,
                                         providers=["CPUExecutionProvider"])
        # The onnxruntime session is thread-safe and can be shared; the lock keeps
        # queueing behavior consistent across threads
        self._lock = threading.Lock()

    def preprocess(self, image_bytes: bytes) -> np.ndarray:
        # Mirrors the training eval_tf step by step: ToTensor + Normalize
        img = Image.open(BytesIO(image_bytes)).convert("RGB").resize((32, 32))
        arr = np.asarray(img, dtype=np.float32) / 255.0          # ToTensor
        arr = arr.transpose(2, 0, 1)                              # HWC -> CHW
        arr = (arr - self.mean) / self.std                        # Normalize
        return arr[None, ...]                                     # (1,3,32,32)

    def predict(self, image_bytes: bytes) -> dict:
        x = self.preprocess(image_bytes)
        with self._lock:
            logits = self.sess.run(["logits"], {"input": x})[0]
        probs = np.exp(logits - logits.max(axis=1, keepdims=True))
        probs = probs / probs.sum(axis=1, keepdims=True)          # softmax
        top = np.argsort(probs[0])[::-1][:5]
        return {"top5": [{"label": self.labels[i], "prob": float(probs[0][i])}
                         for i in top]}

Why preprocessing lives on the server

Training-time augmentation (RandomCrop and friends) belongs only in the training data pipeline, but ToTensor + Normalize is part of the input distribution and must be reproduced server-side. Leaving preprocessing to the client means "the distribution the model sees no longer matches training" — the #1 pitfall in Common Pitfalls and Anti-Patterns.

The matching unit test (3 lines as a placeholder, about 20 in real life):

python
# tests/test_inference.py
def test_preprocess_matches_training():
    # Use an image from the training set: hand-compute ToTensor+Normalize and
    # compare against preprocess(). Assert np.allclose(...) — this is the
    # automated acceptance test for pre/post-processing consistency
    ...

5. Step 4: Serving with FastAPI ​

The service layer does exactly three things: accept requests, call the inference class, return a structured response. Business logic (retries, authentication, rate limiting) does not belong here — that is the gateway's job (covered in Model Gateway and Canary Releases). For a deeper FastAPI walkthrough see An Online Service with FastAPI + Docker.

python
# app/main.py
import time
from fastapi import FastAPI, File, UploadFile, HTTPException
from app.inference import ImageClassifier
from app.metrics import metrics

app = FastAPI(title="image-classifier")
clf = ImageClassifier("models/model.onnx", "models/meta.json")


@app.get("/healthz")
def healthz():
    return {"status": "ok"}          # For K8s/Docker health checks; no inference logic here


@app.post("/predict")
def predict(file: UploadFile = File(...)):
    data = file.file.read()
    t0 = time.perf_counter()
    try:
        result = clf.predict(data)
    except Exception as e:           # Corrupt image, etc. — return 400, not 500
        raise HTTPException(status_code=400, detail=f"bad image: {e}")
    latency_ms = (time.perf_counter() - t0) * 1000
    metrics.observe(latency_ms)      # 7. Instrumentation: counters, histograms
    return {"result": result, "latency_ms": round(latency_ms, 2)}

Launch commands (never run a single uvicorn worker in production — see pitfall 5):

bash
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 8000 --workers 4
curl -F "file=@cat.jpg" http://localhost:8000/predict
curl http://localhost:8000/healthz

The single-worker trap

uvicorn --workers 1 is fine for local debugging, but in production one process means one copy of the model and single-threaded request handling. With --workers 4, every worker loads its own copy of the model (4x the memory). This is the memory/VRAM accounting mistake that trips up deployment newcomers most often (Memory Leaks and OOM has the full analysis).

6. Step 5: Dockerize ​

The goal of containerization is reproduction on any machine with one command. Build the image in two stages — install dependencies first, then copy the model and code — so the model layer and the dependency layer cache separately.

dockerfile
# Dockerfile — built on the slim image, at least half the size
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app/ app/
COPY models/ models/
EXPOSE 8000
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "4"]
yaml
# docker-compose.yml — service + monitoring, one command
services:
  api:
    build: .
    ports: ["8000:8000"]
    restart: unless-stopped
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8000/healthz"]
      interval: 30s
      timeout: 5s
      retries: 3
  prometheus:
    image: prom/prometheus:v2.53.0
    volumes: ["./prometheus.yml:/etc/prometheus/prometheus.yml"]
    ports: ["9090:9090"]

Build and verify:

bash
docker compose up -d --build
docker compose ps            # status should read healthy
curl -F "file=@cat.jpg" http://localhost:8000/predict

Write requirements.txt with pinned versions, never >=:

text
fastapi==0.115.6
uvicorn[standard]==0.32.1
onnxruntime==1.19.2
Pillow==11.0.0
numpy==1.26.4
prometheus-client==0.21.1

7. Step 6: Local Load Testing — Baseline First, Optimization Later ​

Use wrk to grab a rough baseline (5 seconds, 4 threads, 100 connections). POSTing images is awkward with wrk, and adding a GET /predict_path?url= endpoint just for testing is a bad idea — the simpler route: load test /healthz to get the bare service-layer throughput, then benchmark the inference function separately. The more complete approach is to load test a path with a test image embedded; for tools and methodology see Load Testing and Capacity Planning.

bash
# Bare service-layer throughput (no inference)
wrk -t4 -c100 -d10s http://localhost:8000/healthz
# Read Requests/sec and Latency P99 from the output; record as the "service-layer baseline"

# Real inference path: script it with locust and POST the image file
pip install locust
python
# locustfile.py — load test against the real inference path
import io
from PIL import Image
from locust import HttpUser, task, between

class ImageUser(HttpUser):
    wait_time = between(0.05, 0.15)          # simulate a 50-150ms gap between requests
    img_bytes = io.BytesIO()                  # one fixed test image at module level, so load-tester IO never becomes the bottleneck
    Image.new("RGB", (32, 32), (128, 64, 200)).save(img_bytes, format="PNG")

    @task
    def predict(self):
        self.client.post("/predict", files={"file": ("t.png", self.img_bytes.getvalue())})

Three disciplines of load testing (details in Load Testing and Capacity Planning):

  1. The load generator must not be the bottleneck — keep the test image in memory; never read from disk inside the loop;
  2. Warm up for at least 30-60 seconds before recording numbers, skipping the inflated readings from JIT/cache warm-up;
  3. Record the environment: load-generator specs, worker count, concurrency, duration — without these the numbers cannot be reproduced.

8. Step 7: Wiring Up Prometheus Metrics ​

Use prometheus-client to expose the three golden signals: request counts, a latency histogram, and error counts.

python
# app/metrics.py — the single, global metrics module
from prometheus_client import Counter, Histogram

REQUESTS = Counter("predict_requests_total", "Total inference requests",
                   ["model", "version"])
ERRORS = Counter("predict_errors_total", "Failed inference requests",
                 ["model", "version", "error_type"])
LATENCY = Histogram("predict_latency_seconds",
                    "Inference latency (seconds)",
                    buckets=(0.005, 0.01, 0.02, 0.05, 0.1, 0.25, 0.5, 1.0))

def observe(latency_ms: float):
    REQUESTS.labels(model="mobilenetv2", version="v1").inc()
    LATENCY.observe(latency_ms / 1000.0)

Expose /metrics in main.py for Prometheus to scrape:

python
from prometheus_client import generate_latest, CONTENT_TYPE_LATEST
from fastapi import Response

@app.get("/metrics")
def metrics_endpoint():
    return Response(generate_latest(), media_type=CONTENT_TYPE_LATEST)
yaml
# prometheus.yml
scrape_configs:
  - job_name: image-classifier
    static_configs:
      - targets: ["api:8000"]        # use the service name inside the compose network
    scrape_interval: 10s

Once it is up (docker compose up -d prometheus), open http://localhost:9090 and query:

text
# Requests per second (5-minute average)
rate(predict_requests_total[5m])
# P99 latency
histogram_quantile(0.99, sum by (le) (rate(predict_latency_seconds_bucket[5m])))

At this point the minimal deployment loop is complete: you can train, export, serve, containerize, load test, and observe. For full metric design, alerting rules, and Grafana dashboards, see Putting Observability into Practice.

9. The Go-Live Checklist: Do Not Call It "Deployed" Yet ​

Tick every item before going to production; if any answer is "No", hold off:

  • [ ] The /healthz endpoint is ready and the container healthcheck is configured;
  • [ ] Structured logging is wired up (JSON, with a request_id — see Putting Observability into Practice);
  • [ ] /metrics carries the three golden signals and Prometheus scrapes them correctly;
  • [ ] A rollback plan exists: the old image is retained and one command switches back (see Release Strategies: Canary Releases and Rollbacks);
  • [ ] The load test report has conclusions: peak QPS, P99, saturation point;
  • [ ] The model-to-code version mapping is recorded (model registry — see The MLOps Pipeline);
  • [ ] Memory/VRAM ceilings are measured and there is an OOM contingency plan (Common Pitfalls and Anti-Patterns, pitfalls 4 and 5).

10. Common Pitfalls (the Three Most Likely in This Workflow) ​

  1. Preprocessing mismatch: training uses ToTensor+Normalize, the server forgets the normalization → accuracy drops 20 points and nobody can figure out why. Fix: define preprocessing exactly once in train_export.py, write it into meta.json, and have the inference class read it from there.
  2. No validation after export: shipping while ONNX and PyTorch outputs disagree. Fix: append an automatic comparison at the end of the export script, np.allclose(atol=1e-4).
  3. Irreproducible load-test numbers: no environment recorded, no warm-up, load-generator IO as the bottleneck. Fix: put the load-test commands and results into the README (Resume Analysis and Packaging has a template).

11. Where to Go Next: From Skeleton to Production ​

This skeleton is the foundation; four paths lead upward. Pick based on your goal:

DirectionHowReferences
PerformanceINT8 quantization (start with PTQ), switch to a TensorRT engineModel Optimization in Practice, TensorRT Edge Deployment
Multi-model / high throughputMove to Triton, enable dynamic batchingMulti-Model Serving with NVIDIA Triton
LLMsMove to vLLM: PagedAttention + continuous batchingLLM Serving with vLLM, LLM Inference
ScaleK8s + KServe/BentoML, autoscalingChoosing Between Frameworks and Platforms

Checklist ​

  • [ ] All seven steps of the loop completed, each with a deliverable and verification;
  • [ ] Offline baseline accuracy recorded in meta.json so production metrics can be compared against it;
  • [ ] Inference class unit tests pass; preprocessing is identical to training;
  • [ ] The container works with a single docker compose up command;
  • [ ] The load test report includes environment, parameters, and QPS/P99 conclusions;
  • [ ] Golden signals are visible in Grafana and every alert has an owner.

Further Reading ​

References ​