Appearance
Model optimization is a set of engineering techniques for making a model smaller, faster, and cheaper to run without giving up acceptable accuracy. Why start here? Because in the vast majority of model deployment projects, the bottleneck is not a lack of compute — it is the default state of "the model is too big and too slow." A 224×224 MobileNetV2 in FP32 costs 14MB and roughly 6 GFLOPs; quantized to INT8 it is 4x smaller, typically 1.5-2x faster on GPU, and uses 4x less memory. And quite often you do not need to learn anything new — just run the tools in order and pocket a 50% win for free. This page answers three questions: where to start, in what order, and how to know when to stop. The core conclusion: PTQ first; QAT/distillation if accuracy falls short; engine-level optimization only when you need extreme performance — with explicit checkpoints and rollback conditions at every stage. For the theory see Quantization; for the classic papers see Classic Papers on Quantization. Optimization gains must be verified before they go live — pair this with Load Testing and Capacity Planning.
The optimization order in one sentence
90% of the gains come from the first two steps: establish a baseline + PTQ quantization. Everything after that is incremental engineering.
1. Set the Goal Before You Touch Anything: Breaking Down Optimization Targets
Optimization is, at its core, a choice among four goals (or a weighted mix of them), and they pull in opposite directions. Skip this step and the work is guaranteed to be wasted:
| Goal | Key metrics | Typical techniques | Lessons learned the hard way |
|---|---|---|---|
| Size | Model file size in MB, container image size | Quantization (FP32→INT8), pruning | Days spent pruning to save 30MB while accuracy drops 5 points |
| Latency | P99 / per-request time | Quantization, TensorRT, batching tuning | Optimizing P50 while ignoring P99 — long-tail requests still time out |
| Throughput | QPS, tokens per second | Continuous batching, concurrency tuning | Fast per request, but it collapses as soon as concurrency rises |
| Power/cost | Cost per 1,000 inferences, watts | Quantization + low-power hardware selection | Running a small model on a GPU at 300W when a CPU could handle it |
Three hard rules:
- Quantize before pruning: quantization costs almost no engineering effort (minutes); pruning requires retraining. Try quantization first.
- Measure latency and throughput separately: quantization helps single-request latency only modestly, but helps throughput enormously (batches get bigger) — do not let a disappointing single-request number talk you out of quantization.
- The goal is meeting the SLO, not squeezing out the last drop: once P99 < 100ms is met, stop — extra optimization is negative ROI. See the SLO definitions in Performance Metrics.
2. The Optimization Roadmap: A Pipeline with Rollbacks
text
Baseline ──▶ PTQ quantization ──▶ Accuracy eval ──▶ Pass? ──▶ Load-test and ship
│ │ │ │
│ │ └── Fail ───▶ QAT / distillation ──▶ Re-evaluate
│ │ │
│ │ └── Need extreme perf ──▶ TensorRT / engine
│ └── Some ops unsupported ──▶ Mixed precision (FP16/INT8 per segment)
└── Latency misses but accuracy is critical ──▶ Start from pruning/distillationEvery node has a rollback: accuracy drops after quantization → fall back to the FP32 baseline and change strategy; QAT is too expensive to train → go back to distillation. The value of this roadmap is that it turns "optimizing on gut feel" into "decision-making backed by evidence". The recording format for each step appears in the optimization tracking table later in this page.
3. Step 1: Measure the Baseline — Without One, Every Optimization Is Just Noise
Before optimizing, spend 30 minutes on a proper measurement and produce three sets of numbers: latency (P50/P99), peak memory/VRAM, and the accuracy metric. This baseline is the anchor for every comparison that follows.
python
# baseline.py — a standard script for baselines, avoiding hand-timing errors
import time
import torch
import onnxruntime as ort
import numpy as np
def measure_ort(onnx_path: str, n: int = 200, warmup: int = 20):
"""Measures the ONNX Runtime latency distribution (in ms)."""
sess = ort.InferenceSession(onnx_path, providers=["CPUExecutionProvider"])
x = np.random.rand(1, 3, 32, 32).astype(np.float32)
for _ in range(warmup): # warm-up, skipping JIT/cache ramp-up
sess.run(None, {"input": x})
lats = []
for _ in range(n):
t0 = time.perf_counter()
sess.run(None, {"input": x})
lats.append((time.perf_counter() - t0) * 1000)
lats.sort()
return {"p50": lats[n // 2], "p99": lats[int(n * 0.99) - 1], "mean": sum(lats) / n}Baseline recording template (save it as perf_baseline.md and commit it with the repo):
| Item | Value | Example |
|---|---|---|
| Environment | CPU/GPU model, driver, engine version | Intel i7-12700 / ONNX Runtime 1.19.2 CPU |
| Input spec | shape, batch, dtype | (1,3,32,32) FP32 |
| Latency P50/P99 | ms | 8.2 / 14.5 |
| Peak memory/VRAM | MB | 220 MB |
| Model size | MB | 14.0 |
| Offline accuracy | Validation-set metric | accuracy 0.812 |
| Thread/batch config | threads, batch | 4 threads, batch=1 |
Never substitute "seems fast" for measurement
Wrapping a single inference call in time.time() can be off by ±30% — no warm-up, insufficient clock resolution. Only the script form above (warm-up + many samples + percentiles) produces numbers you can actually compare.
4. Step 2: PTQ Quantization in Practice (Try This First)
PTQ (post-training quantization) needs no retraining and is the highest-ROI step. Two mainstream entry points: ONNX Runtime's quantize_static (fast on CPU) and PyTorch's torch.ao.quantization.
4.1 Static Quantization and Calibration Data
Static quantization needs a small batch of "calibration data" to determine the numeric range of each tensor (scale/zero_point). Calibration data must come from the real input distribution — calibrating on random noise can double the quantization error out of nowhere, and it is the detail newcomers get wrong most often.
python
# quantize_ptq.py — ONNX Runtime static quantization
import onnx
from onnxruntime.quantization import quantize_static, CalibrationMethod
from onnxruntime.quantization.calibrate import CalibrationDataReader
class CalibReader(CalibrationDataReader):
"""Takes 200 images from the validation set as calibration data; the input name must match the export."""
def __init__(self, dataloader, input_name="input"):
self.iter = iter(dataloader)
self.input_name = input_name
def get_next(self):
try:
x, _ = next(self.iter)
return {self.input_name: x.numpy()}
except StopIteration:
return None
# 200 calibration images covering the class distribution; more is not better — 256 is plenty
calib = CalibReader(DataLoader(ds_eval, batch_size=32)[:200])
quantize_static(
"models/model.onnx", "models/model_int8.onnx",
calib, per_channel=True,
calibration_method=CalibrationMethod.MinMax, # MinMax is the default; ENTROPY is steadier for distribution-sensitive models
)
print("INT8 model exported; size should be about 1/4 of FP32")4.2 Do Three Things Immediately After Quantizing
bash
# 1) Size: should be ≈ 25% of FP32 (14MB -> 3.5MB)
ls -lh models/model.onnx models/model_int8.onnx
# 2) Accuracy: run the validation set once and compare against the baseline table
python evaluate.py --model models/model_int8.onnx # expect an accuracy drop < 1%
# 3) Latency/memory: rerun baseline.py and compare against the baseline table
python baseline.py --model models/model_int8.onnxWhat to expect from quantization (rules of thumb)
- Size: INT8 ≈ 1/4 of FP32; FP16 ≈ 1/2;
- Latency: on CPU, INT8 is typically 1.5-2x faster; on GPU it depends on whether the ops are covered by accelerated kernels;
- Accuracy: classification/detection models usually lose 0.5-2% after PTQ. If you lose more than 5%, check the calibration data first instead of rushing into QAT.
5. Step 3: Accuracy Evaluation and the Rollback Decision
After quantization, exactly one question matters: is this accuracy loss acceptable? The test is not "how far from FP32" but "did it cross a business threshold?"
| Scenario | Verdict | Action |
|---|---|---|
| accuracy drop ≤ 1% | Pass | Ship INT8 and move on to load-test verification |
| Drop of 1%-5% | Borderline | Try a different calibration method (MinMax→Entropy) and different calibration data first; escalate to QAT only if that fails |
| Drop > 5% | Fail | Roll back to FP32 and take the distillation or pruning route |
python
# evaluate.py — the automated version of the accuracy comparison, with structured output
import onnxruntime as ort
import numpy as np
def eval_onnx(path, loader):
sess = ort.InferenceSession(path, providers=["CPUExecutionProvider"])
correct = total = 0
for x, y in loader:
pred = sess.run(None, {"input": x.numpy()})[0].argmax(1)
correct += (pred == y.numpy()).sum().item()
total += y.size(0)
return correct / total
print("FP32:", eval_onnx("models/model.onnx", loader))
print("INT8:", eval_onnx("models/model_int8.onnx", loader))Rollback is a normal part of the process, not a failure. Bake the evaluation script into CI so that every quantized build is automatically compared against the baseline — this is also the seed of the "model validation" stage in The MLOps Pipeline.
6. Step 4: QAT and Distillation (Two Paths When Accuracy Falls Short)
When PTQ misses the bar, there are two main lines of attack:
6.1 QAT (Quantization-Aware Training)
QAT simulates quantization error inside the training process (fake-quant ops) so that the weights learn to preserve accuracy after quantization. The official PyTorch path (torch.ao.quantization):
python
# qat.py — key flow only; see the PyTorch docs for complete code
import torch
from torch.ao.quantization import (
get_default_qat_qconfig_mapping, QATQuantizer, prepare_qat, convert)
qat = QATQuantizer(model) # or prepare_qat(model, qconfig) manually
qat.qconfig = get_default_qat_qconfig_mapping("x86") # pick the qconfig for your deployment backend
model_qat = prepare_qat(model, inplace=False)
# Fine-tune for 3-5 more epochs on the training set (drop the learning rate to about 1/10)
# ...
model_int8 = convert(model_qat) # convert to a genuinely quantized model once training is doneQAT's cost is training time plus training data and a GPU; its benefit is holding on to 1-3% more accuracy than PTQ. Rule of thumb: QAT is worth it when PTQ loses more than 3%; at a 1% loss, QAT's ROI is poor.
6.2 Distillation (Knowledge Distillation)
Use a large model (the teacher) and its soft labels to teach a small model (the student), so the smaller model approaches the teacher's accuracy at a fraction of the size. Distillation and quantization are orthogonal: you can distill into a smaller FP32 model and then quantize that to INT8. For theory and tools see Model Compression. Typical payoff: at the same 80% accuracy, the distilled model is another order of magnitude smaller (a ResNet-50-class teacher distilling down to a MobileNet-class student).
python
# distill.py — key code: the loss consumes both hard labels and teacher soft labels
alpha, T = 0.7, 3.0 # soft-label weight 0.7, temperature 3
kl = nn.KLDivLoss(reduction="batchmean")
with torch.no_grad():
t_logits = teacher(x)
student_loss = F.cross_entropy(s_logits, y)
distill_loss = kl(F.log_softmax(s_logits / T, dim=1),
F.softmax(t_logits / T, dim=1))
loss = alpha * T * T * distill_loss + (1 - alpha) * student_lossDo not start distilling before you define the target
Before distilling, ask: what are the student's deployment constraints (how much memory? which engine? what latency budget?). Fix the constraints first, then pick the teacher and student — otherwise you will finish distilling only to discover that the student's advantage does not hold on the target hardware, and a week is gone.
7. Step 5: Engine-Level Optimization (TensorRT and Friends)
Quantization slims the model; the engine accelerates execution: layer fusion, kernel selection, memory planning. TensorRT is typically another 30-60% faster than ONNX Runtime on NVIDIA GPUs, and INT8 + engine is the common end state on GPU. For the full workflow (the trtexec CLI, engine serialization, hardware binding) see TensorRT Edge Deployment.
bash
# Build an FP16 engine from ONNX with trtexec and measure throughput
trtexec --onnx=models/model.onnx \
--saveEngine=models/model_fp16.engine \
--fp16 \
--shapes=input:1x3x32x32 \
--avgRuns=100Three preconditions for engine-level optimization:
- Quantize before you build the engine: TensorRT can do INT8 internally, but a model quantized externally is more controllable;
- Engines are bound to hardware: an
.enginefile is tightly coupled to the GPU model and TensorRT version; switching cards means rebuilding — the direct consequence of "the engine is a non-portable format" from Model Formats; - An engine only makes the compute fast; it does not keep the pipeline fed: input pipeline overhead (decoding, preprocessing) becomes the new CPU-side bottleneck. Watch end-to-end P99 in load tests, not just the engine-internal time.
8. Per-Step Checklist (What to Record)
| Step | Must record | Commonly missed |
|---|---|---|
| Baseline | Environment, P50/P99, memory, accuracy, model size | Forgetting to record batch size and thread count |
| PTQ quantization | Calibration data source and count, calibration method, per-channel or not | Calibrating on random noise |
| Accuracy evaluation | FP32 vs INT8 comparison, business threshold | Looking only at the relative delta, not whether the business floor was broken |
| QAT/distillation | Training epochs, learning rate, teacher configuration | Not recording training cost (GPU-hours) |
| Engine optimization | Engine version, GPU model, build flags | Not recording the engine-to-hardware binding |
9. Optimization Tracking Table Template
Aggregate the whole pipeline's numbers into one table that travels with each model version (it maps to the model registry entry in The MLOps Pipeline):
| Version | Approach | Size | P50(ms) | P99(ms) | Mem(MB) | Accuracy | Decision |
|---|---|---|---|---|---|---|---|
| v1 | FP32 baseline | 14.0 | 8.2 | 14.5 | 220 | 0.812 | Baseline |
| v2 | INT8 (MinMax) | 3.6 | 4.8 | 8.9 | 96 | 0.806 | Ship |
| v3 | INT8 + TensorRT | 3.6 | 2.9 | 5.4 | 88 | 0.806 | Ship |
Decision rule: only candidates that are both "within the business accuracy threshold" and "meeting the latency/throughput SLO" enter the ship queue; if either condition fails, roll back to the previous row.
10. Decision Table: Which Optimization for Which Situation
| Your situation | Recommended path | Why |
|---|---|---|
| First optimization, tight deadline | PTQ quantization → load test | No retraining; results within hours |
| Accuracy drops > 3% after PTQ | QAT (if you have training resources) | Keeps 1-3% more accuracy than PTQ |
| Size is a hard constraint (edge/mobile) | Distillation → quantization | Another order of magnitude smaller |
| Accuracy-critical and both PTQ/QAT fail | Pruning + retraining | Structural slimming; the most expensive option, so it goes last |
| NVIDIA GPU, latency-sensitive | Quantization + TensorRT | Another 30-60% from the engine |
| LLM workloads | GPTQ/AWQ + vLLM | Weight quantization is the standard for LLMs — see LLM Inference |
Why pruning comes last
Structured pruning (structured sparsity) requires retraining, the tooling is less mature, and its gains are often subsumed by quantization. Quantize first, prune later is the community consensus; treat pruning as the fallback for when quantization cannot meet a size constraint, not as the starting point.
Trade-offs
Optimization trades accuracy and engineering effort for size, latency, and cost. Three trade-offs that are easy to overlook: calibration data is part of your training data, so it touches data compliance; engines bind to hardware, which weakens portability (switch cloud instance types and you rebuild); quantized models are hard to fine-tune further, so every model iteration has to re-run the quantization pipeline — which is exactly why the pipeline itself should be scripted and wired into CI, not operated by hand every time.
Further Reading
- Quantization — INT8 fundamentals, scale/zero_point, and the theory behind PTQ/QAT
- Classic Papers on Quantization — close readings of GPTQ, AWQ, and similar work: why these calibration schemes look the way they do
- Model Compression — the complete methodology for distillation and pruning
- TensorRT Edge Deployment — from ONNX to engine, end to end
- Load Testing and Capacity Planning — how to accept optimization gains with load tests
- Deploy a Model from Scratch — plug the optimized model into the full deployment loop
References
- ONNX Runtime quantization docs: https://onnxruntime.ai/docs/performance/model-optimizations/quantization.html
- PyTorch quantization docs (PTQ/QAT): https://pytorch.org/docs/stable/quantization.html
- NVIDIA TensorRT docs: https://docs.nvidia.com/deeplearning/tensorrt/
- Hugging Face Optimum (ONNX export and quantization): https://huggingface.co/docs/optimum/index