Skip to content

Model Optimization in Practice

At a glance How do you slim down and speed up a model before it ships? This page gives a hands-on roadmap — quantize first (start with PTQ), move to QAT/distillation if accuracy falls short, prune when needed — with tools and checkpoints at every step, closing with a decision table.

Model optimization is a set of engineering techniques for making a model smaller, faster, and cheaper to run without giving up acceptable accuracy. Why start here? Because in the vast majority of model deployment projects, the bottleneck is not a lack of compute — it is the default state of "the model is too big and too slow." A 224×224 MobileNetV2 in FP32 costs 14MB and roughly 6 GFLOPs; quantized to INT8 it is 4x smaller, typically 1.5-2x faster on GPU, and uses 4x less memory. And quite often you do not need to learn anything new — just run the tools in order and pocket a 50% win for free. This page answers three questions: where to start, in what order, and how to know when to stop. The core conclusion: PTQ first; QAT/distillation if accuracy falls short; engine-level optimization only when you need extreme performance — with explicit checkpoints and rollback conditions at every stage. For the theory see Quantization; for the classic papers see Classic Papers on Quantization. Optimization gains must be verified before they go live — pair this with Load Testing and Capacity Planning.

The optimization order in one sentence

90% of the gains come from the first two steps: establish a baseline + PTQ quantization. Everything after that is incremental engineering.

1. Set the Goal Before You Touch Anything: Breaking Down Optimization Targets ​

Optimization is, at its core, a choice among four goals (or a weighted mix of them), and they pull in opposite directions. Skip this step and the work is guaranteed to be wasted:

GoalKey metricsTypical techniquesLessons learned the hard way
SizeModel file size in MB, container image sizeQuantization (FP32→INT8), pruningDays spent pruning to save 30MB while accuracy drops 5 points
LatencyP99 / per-request timeQuantization, TensorRT, batching tuningOptimizing P50 while ignoring P99 — long-tail requests still time out
ThroughputQPS, tokens per secondContinuous batching, concurrency tuningFast per request, but it collapses as soon as concurrency rises
Power/costCost per 1,000 inferences, wattsQuantization + low-power hardware selectionRunning a small model on a GPU at 300W when a CPU could handle it

Three hard rules:

  1. Quantize before pruning: quantization costs almost no engineering effort (minutes); pruning requires retraining. Try quantization first.
  2. Measure latency and throughput separately: quantization helps single-request latency only modestly, but helps throughput enormously (batches get bigger) — do not let a disappointing single-request number talk you out of quantization.
  3. The goal is meeting the SLO, not squeezing out the last drop: once P99 < 100ms is met, stop — extra optimization is negative ROI. See the SLO definitions in Performance Metrics.

2. The Optimization Roadmap: A Pipeline with Rollbacks ​

text
  Baseline ──▶ PTQ quantization ──▶ Accuracy eval ──▶ Pass? ──▶ Load-test and ship
    │               │                  │             │
    │               │                  └── Fail ───▶ QAT / distillation ──▶ Re-evaluate
    │               │                                              │
    │               │                                              └── Need extreme perf ──▶ TensorRT / engine
    │               └── Some ops unsupported ──▶ Mixed precision (FP16/INT8 per segment)
    └── Latency misses but accuracy is critical ──▶ Start from pruning/distillation

Every node has a rollback: accuracy drops after quantization → fall back to the FP32 baseline and change strategy; QAT is too expensive to train → go back to distillation. The value of this roadmap is that it turns "optimizing on gut feel" into "decision-making backed by evidence". The recording format for each step appears in the optimization tracking table later in this page.

3. Step 1: Measure the Baseline — Without One, Every Optimization Is Just Noise ​

Before optimizing, spend 30 minutes on a proper measurement and produce three sets of numbers: latency (P50/P99), peak memory/VRAM, and the accuracy metric. This baseline is the anchor for every comparison that follows.

python
# baseline.py — a standard script for baselines, avoiding hand-timing errors
import time
import torch
import onnxruntime as ort
import numpy as np

def measure_ort(onnx_path: str, n: int = 200, warmup: int = 20):
    """Measures the ONNX Runtime latency distribution (in ms)."""
    sess = ort.InferenceSession(onnx_path, providers=["CPUExecutionProvider"])
    x = np.random.rand(1, 3, 32, 32).astype(np.float32)
    for _ in range(warmup):                     # warm-up, skipping JIT/cache ramp-up
        sess.run(None, {"input": x})
    lats = []
    for _ in range(n):
        t0 = time.perf_counter()
        sess.run(None, {"input": x})
        lats.append((time.perf_counter() - t0) * 1000)
    lats.sort()
    return {"p50": lats[n // 2], "p99": lats[int(n * 0.99) - 1], "mean": sum(lats) / n}

Baseline recording template (save it as perf_baseline.md and commit it with the repo):

ItemValueExample
EnvironmentCPU/GPU model, driver, engine versionIntel i7-12700 / ONNX Runtime 1.19.2 CPU
Input specshape, batch, dtype(1,3,32,32) FP32
Latency P50/P99ms8.2 / 14.5
Peak memory/VRAMMB220 MB
Model sizeMB14.0
Offline accuracyValidation-set metricaccuracy 0.812
Thread/batch configthreads, batch4 threads, batch=1

Never substitute "seems fast" for measurement

Wrapping a single inference call in time.time() can be off by ±30% — no warm-up, insufficient clock resolution. Only the script form above (warm-up + many samples + percentiles) produces numbers you can actually compare.

4. Step 2: PTQ Quantization in Practice (Try This First) ​

PTQ (post-training quantization) needs no retraining and is the highest-ROI step. Two mainstream entry points: ONNX Runtime's quantize_static (fast on CPU) and PyTorch's torch.ao.quantization.

4.1 Static Quantization and Calibration Data ​

Static quantization needs a small batch of "calibration data" to determine the numeric range of each tensor (scale/zero_point). Calibration data must come from the real input distribution — calibrating on random noise can double the quantization error out of nowhere, and it is the detail newcomers get wrong most often.

python
# quantize_ptq.py — ONNX Runtime static quantization
import onnx
from onnxruntime.quantization import quantize_static, CalibrationMethod
from onnxruntime.quantization.calibrate import CalibrationDataReader

class CalibReader(CalibrationDataReader):
    """Takes 200 images from the validation set as calibration data; the input name must match the export."""
    def __init__(self, dataloader, input_name="input"):
        self.iter = iter(dataloader)
        self.input_name = input_name
    def get_next(self):
        try:
            x, _ = next(self.iter)
            return {self.input_name: x.numpy()}
        except StopIteration:
            return None

# 200 calibration images covering the class distribution; more is not better — 256 is plenty
calib = CalibReader(DataLoader(ds_eval, batch_size=32)[:200])

quantize_static(
    "models/model.onnx", "models/model_int8.onnx",
    calib, per_channel=True,
    calibration_method=CalibrationMethod.MinMax,   # MinMax is the default; ENTROPY is steadier for distribution-sensitive models
)
print("INT8 model exported; size should be about 1/4 of FP32")

4.2 Do Three Things Immediately After Quantizing ​

bash
# 1) Size: should be ≈ 25% of FP32 (14MB -> 3.5MB)
ls -lh models/model.onnx models/model_int8.onnx
# 2) Accuracy: run the validation set once and compare against the baseline table
python evaluate.py --model models/model_int8.onnx   # expect an accuracy drop < 1%
# 3) Latency/memory: rerun baseline.py and compare against the baseline table
python baseline.py --model models/model_int8.onnx

What to expect from quantization (rules of thumb)

  • Size: INT8 ≈ 1/4 of FP32; FP16 ≈ 1/2;
  • Latency: on CPU, INT8 is typically 1.5-2x faster; on GPU it depends on whether the ops are covered by accelerated kernels;
  • Accuracy: classification/detection models usually lose 0.5-2% after PTQ. If you lose more than 5%, check the calibration data first instead of rushing into QAT.

5. Step 3: Accuracy Evaluation and the Rollback Decision ​

After quantization, exactly one question matters: is this accuracy loss acceptable? The test is not "how far from FP32" but "did it cross a business threshold?"

ScenarioVerdictAction
accuracy drop ≤ 1%PassShip INT8 and move on to load-test verification
Drop of 1%-5%BorderlineTry a different calibration method (MinMax→Entropy) and different calibration data first; escalate to QAT only if that fails
Drop > 5%FailRoll back to FP32 and take the distillation or pruning route
python
# evaluate.py — the automated version of the accuracy comparison, with structured output
import onnxruntime as ort
import numpy as np

def eval_onnx(path, loader):
    sess = ort.InferenceSession(path, providers=["CPUExecutionProvider"])
    correct = total = 0
    for x, y in loader:
        pred = sess.run(None, {"input": x.numpy()})[0].argmax(1)
        correct += (pred == y.numpy()).sum().item()
        total += y.size(0)
    return correct / total

print("FP32:", eval_onnx("models/model.onnx", loader))
print("INT8:", eval_onnx("models/model_int8.onnx", loader))

Rollback is a normal part of the process, not a failure. Bake the evaluation script into CI so that every quantized build is automatically compared against the baseline — this is also the seed of the "model validation" stage in The MLOps Pipeline.

6. Step 4: QAT and Distillation (Two Paths When Accuracy Falls Short) ​

When PTQ misses the bar, there are two main lines of attack:

6.1 QAT (Quantization-Aware Training) ​

QAT simulates quantization error inside the training process (fake-quant ops) so that the weights learn to preserve accuracy after quantization. The official PyTorch path (torch.ao.quantization):

python
# qat.py — key flow only; see the PyTorch docs for complete code
import torch
from torch.ao.quantization import (
    get_default_qat_qconfig_mapping, QATQuantizer, prepare_qat, convert)

qat = QATQuantizer(model)                       # or prepare_qat(model, qconfig) manually
qat.qconfig = get_default_qat_qconfig_mapping("x86")   # pick the qconfig for your deployment backend
model_qat = prepare_qat(model, inplace=False)
# Fine-tune for 3-5 more epochs on the training set (drop the learning rate to about 1/10)
# ...
model_int8 = convert(model_qat)                 # convert to a genuinely quantized model once training is done

QAT's cost is training time plus training data and a GPU; its benefit is holding on to 1-3% more accuracy than PTQ. Rule of thumb: QAT is worth it when PTQ loses more than 3%; at a 1% loss, QAT's ROI is poor.

6.2 Distillation (Knowledge Distillation) ​

Use a large model (the teacher) and its soft labels to teach a small model (the student), so the smaller model approaches the teacher's accuracy at a fraction of the size. Distillation and quantization are orthogonal: you can distill into a smaller FP32 model and then quantize that to INT8. For theory and tools see Model Compression. Typical payoff: at the same 80% accuracy, the distilled model is another order of magnitude smaller (a ResNet-50-class teacher distilling down to a MobileNet-class student).

python
# distill.py — key code: the loss consumes both hard labels and teacher soft labels
alpha, T = 0.7, 3.0          # soft-label weight 0.7, temperature 3
kl = nn.KLDivLoss(reduction="batchmean")
with torch.no_grad():
    t_logits = teacher(x)
student_loss = F.cross_entropy(s_logits, y)
distill_loss = kl(F.log_softmax(s_logits / T, dim=1),
                  F.softmax(t_logits / T, dim=1))
loss = alpha * T * T * distill_loss + (1 - alpha) * student_loss

Do not start distilling before you define the target

Before distilling, ask: what are the student's deployment constraints (how much memory? which engine? what latency budget?). Fix the constraints first, then pick the teacher and student — otherwise you will finish distilling only to discover that the student's advantage does not hold on the target hardware, and a week is gone.

7. Step 5: Engine-Level Optimization (TensorRT and Friends) ​

Quantization slims the model; the engine accelerates execution: layer fusion, kernel selection, memory planning. TensorRT is typically another 30-60% faster than ONNX Runtime on NVIDIA GPUs, and INT8 + engine is the common end state on GPU. For the full workflow (the trtexec CLI, engine serialization, hardware binding) see TensorRT Edge Deployment.

bash
# Build an FP16 engine from ONNX with trtexec and measure throughput
trtexec --onnx=models/model.onnx \
        --saveEngine=models/model_fp16.engine \
        --fp16 \
        --shapes=input:1x3x32x32 \
        --avgRuns=100

Three preconditions for engine-level optimization:

  1. Quantize before you build the engine: TensorRT can do INT8 internally, but a model quantized externally is more controllable;
  2. Engines are bound to hardware: an .engine file is tightly coupled to the GPU model and TensorRT version; switching cards means rebuilding — the direct consequence of "the engine is a non-portable format" from Model Formats;
  3. An engine only makes the compute fast; it does not keep the pipeline fed: input pipeline overhead (decoding, preprocessing) becomes the new CPU-side bottleneck. Watch end-to-end P99 in load tests, not just the engine-internal time.

8. Per-Step Checklist (What to Record) ​

StepMust recordCommonly missed
BaselineEnvironment, P50/P99, memory, accuracy, model sizeForgetting to record batch size and thread count
PTQ quantizationCalibration data source and count, calibration method, per-channel or notCalibrating on random noise
Accuracy evaluationFP32 vs INT8 comparison, business thresholdLooking only at the relative delta, not whether the business floor was broken
QAT/distillationTraining epochs, learning rate, teacher configurationNot recording training cost (GPU-hours)
Engine optimizationEngine version, GPU model, build flagsNot recording the engine-to-hardware binding

9. Optimization Tracking Table Template ​

Aggregate the whole pipeline's numbers into one table that travels with each model version (it maps to the model registry entry in The MLOps Pipeline):

VersionApproachSizeP50(ms)P99(ms)Mem(MB)AccuracyDecision
v1FP32 baseline14.08.214.52200.812Baseline
v2INT8 (MinMax)3.64.88.9960.806Ship
v3INT8 + TensorRT3.62.95.4880.806Ship

Decision rule: only candidates that are both "within the business accuracy threshold" and "meeting the latency/throughput SLO" enter the ship queue; if either condition fails, roll back to the previous row.

10. Decision Table: Which Optimization for Which Situation ​

Your situationRecommended pathWhy
First optimization, tight deadlinePTQ quantization → load testNo retraining; results within hours
Accuracy drops > 3% after PTQQAT (if you have training resources)Keeps 1-3% more accuracy than PTQ
Size is a hard constraint (edge/mobile)Distillation → quantizationAnother order of magnitude smaller
Accuracy-critical and both PTQ/QAT failPruning + retrainingStructural slimming; the most expensive option, so it goes last
NVIDIA GPU, latency-sensitiveQuantization + TensorRTAnother 30-60% from the engine
LLM workloadsGPTQ/AWQ + vLLMWeight quantization is the standard for LLMs — see LLM Inference

Why pruning comes last

Structured pruning (structured sparsity) requires retraining, the tooling is less mature, and its gains are often subsumed by quantization. Quantize first, prune later is the community consensus; treat pruning as the fallback for when quantization cannot meet a size constraint, not as the starting point.

Trade-offs ​

Optimization trades accuracy and engineering effort for size, latency, and cost. Three trade-offs that are easy to overlook: calibration data is part of your training data, so it touches data compliance; engines bind to hardware, which weakens portability (switch cloud instance types and you rebuild); quantized models are hard to fine-tune further, so every model iteration has to re-run the quantization pipeline — which is exactly why the pipeline itself should be scripted and wired into CI, not operated by hand every time.

Further Reading ​

References ​