Appearance
Model Formats and Conversion
The one-sentence definition: a model format is the serialization container for weights and the computation graph—to run efficient inference on the target hardware, the model first has to be "translated" into a language that hardware understands, and that translation is model conversion.
Industry insight: the model format is the most underrated link in the deployment chain. The .pt file PyTorch produces carries Python runtime dependencies; forget one dynamic_axes entry when exporting to ONNX and dynamic batching fails outright in production; forget to re-export after quantization and your INT8 speedup evaporates. Industry experience: a sizable share of deployment incidents aren't "the model is bad"—they're "wrong format, bad export, missing validation." Understanding the format ecosystem is laying the foundation for the entire deployment chain.
1. Why an Intermediate Format Is Needed
1.1 Three Decouplings
- Decoupling from the training framework: PyTorch is great for iterating during training, but production wants no Python or framework dependencies—you can't drag a 2GB PyTorch environment around just to run one inference on a 500MB model.
- Decoupling from hardware: the same model has to run on NVIDIA GPUs (TensorRT), Intel/AMD CPUs (OpenVINO/ONNX Runtime), phones (TFLite/Core ML), and browsers (WebAssembly)—you can't make every hardware target integrate with PyTorch directly.
- Decoupling from the inference runtime: ONNX is a standard intermediate representation (IR) that ONNX Runtime, Triton, and TensorRT can all consume, avoiding vendor lock-in.
1.2 Two Conversion Routes
text
Route 1 (the standard route):
PyTorch .pt ──export──► ONNX ──optimize/compile──► TensorRT .engine / OpenVINO IR
(developer-friendly, portable) (specialized to the target hardware, peak performance)
Route 2 (the direct route):
PyTorch .pt ──torch.compile / direct load──► PyTorch runtime
(skips the intermediate format, sacrifices some performance or portability)Engineering practice favors the standard route: with ONNX as a stable intermediate layer, you can validate correctness quickly during development and then recompile for each target hardware. In the LLM era, GGUF is a special case of "the format is the deployment"—see Section 6 of this chapter.
2. The Mainstream Format Landscape
| Format | Ecosystem | Target hardware | One-line summary |
|---|---|---|---|
.pt / .pth | PyTorch | General (needs PyTorch runtime) | Training artifact, holds weights and structure (or a state_dict) |
| TorchScript | PyTorch | General (runs in C++ too) | PyTorch's serialized graph format, no external dependencies |
| ONNX | Open standard | Cross-framework, cross-hardware | The de facto interchange format, broadest ecosystem |
.engine | NVIDIA TensorRT | NVIDIA GPUs | Compiled artifact, peak performance, hardware-bound |
.tflite | TensorFlow Lite | Mobile/edge/microcontrollers | De facto mobile standard, supports quantization |
| Core ML | Apple | Apple silicon | Apple ecosystem's native format |
| OpenVINO IR | Intel | Intel CPU/GPU/NPU | Optimized for Intel hardware, a strong CPU inference tool |
| GGUF | llama.cpp ecosystem | Both CPU and GPU | LLM-specific, embeds quantization schemes and tokenizer |
| SafeTensors | Hugging Face | General | Fast, safe weight container, no code execution |
Remember Three Watersheds
- ONNX: the portability watershed—use it and your model isn't tied to any framework;
- TensorRT
.engine/ OpenVINO IR: the performance watershed—use it and you're bound to one specific piece of hardware; - GGUF / TFLite: the scenario watershed—GGUF for running LLMs locally, TFLite for running on-device.
SafeTensors: Why the LLM Community Is Abandoning .bin
Hugging Face's pytorch_model.bin is essentially pickle, which executes arbitrary code on load—download a malicious weight file and you're poisoned. SafeTensors is a pure data format (JSON header + binary tensors): fast to load (zero deserialization overhead; loading 7B weights takes seconds) and safe. It's now the recommended container for LLM weight distribution. For the related supply-chain risks, see Security, Privacy, and Compliance.
3. The Standard Conversion Pipeline
text
Training output .pt
│ ① export
▼
ONNX graph (with dynamic axes, opset version)
│ ② optimize graph optimization / precision calibration (INT8 needs calibration)
▼
Optimized ONNX
│ ③ build targeting the destination hardware
▼
TensorRT engine / OpenVINO IR / TFLite
│ ④ validate numerical comparison + measured performance
▼
ProductionEvery step can introduce problems—here they are one by one.
4. ONNX Export in Detail
4.1 Using torch.onnx.export Correctly
python
import torch
import torch.onnx
model = MyModel().eval().cuda()
dummy = torch.randn(1, 3, 224, 224).cuda() # input shape must match production
torch.onnx.export(
model,
dummy,
"model.onnx",
input_names=["input"], # name them; makes backend binding easier
output_names=["logits"],
dynamic_axes={ # critical: declare dynamic axes
"input": {0: "batch", 2: "height", 3: "width"},
"logits": {0: "batch"},
},
opset_version=17, # opset version, see below
do_constant_folding=True, # constant folding, on by default
)Three frequent mistakes:
- Forgetting
eval(): BatchNorm/Dropout behave differently in training mode, so the exported graph is wrong. - Missing
dynamic_axes: shapes get frozen todummy's by default, and batch=8 in production fails outright. Declare every axis that might vary. - Export with realistic input shapes: shape is a hard constraint at export time; a mismatch with the real input leads to a broken graph.
4.2 What the Opset Version Is
The ONNX spec evolves operator semantics version by version; opset_version=17 means you're using the 17th operator set. Rule of thumb: use the highest version that covers all your operators and is supported by the target runtime. ONNX Runtime maintains an official operator compatibility matrix—for example, some new semantics from opset 18 aren't recognized by older runtime versions. The practical workflow:
text
Check the runtime's docs → pick the highest opset it supports → on "unsupported operator" errors, lower the version or rewrite the graph4.3 Validate After Export with onnx.checker and onnxruntime
python
import onnx
import onnxruntime as ort
onnx.checker.check_model("model.onnx") # structural validity
sess = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
# feed a few real inputs and compare ONNX outputs vs PyTorch outputs (see Section 6 for numerical validation)5. Operator Incompatibility: The Most Common Deployment Trap
5.1 What the Problem Looks Like
Export to ONNX fails, or building on the target engine reports "Unsupported operator / op type not registered." Common causes:
- Dynamic control flow:
if/forthat depend on tensor values—ONNX doesn't support this; rewrite usingtorch.whereor fixed loops. - Training-only operators:
F.dropout(fixed byevalmode), anythingtorch.autograd-related. - Custom / new operators: e.g., older versions don't support
torch.roll, and some attention implementations. - Unsupported dtypes: some engines reject FP64 or
int64inputs.
5.2 Three Mitigation Strategies (Cheapest First)
| Strategy | How | Cost | When to use |
|---|---|---|---|
| Replace the operator | Rewrite with equivalent standard ops (e.g., attention as matmul+softmax) | Low | Most cases |
| Partition the graph | Run unsupported subgraphs as Python fallback; everything else on the engine | Medium | A few stubborn operators |
| Fall back to CPU | Force that operator to run on CPU | Low | Rare operators, ultra-low-traffic paths |
The Cost of Partitioning and Fallbacks
Once the graph is split, intermediate tensors get copied back and forth, and performance can drop 20%+. Replace operators first; partition only as a last resort; reserve CPU fallback for paths that "run a handful of times a year."
5.3 The Real Fix: The Newer the Operators, the Earlier You Validate
New models (especially novel Transformer architectures) should get an export smoke test during training—not a week before launch, when you discover the engine doesn't support some attention operator. This belongs to the "model deployability check" gate in the MLOps Deployment Pipeline.
6. Post-Conversion Numerical Validation
6.1 Why Validation Is Non-Negotiable
Export/compilation doesn't change the math, but operator reordering, low-precision kernels, and different accumulation orders all introduce numerical differences. If you skip validation and ship, then when business metrics go wrong you can't tell whether it's the model or the conversion.
6.2 Three-Layer Validation
python
import numpy as np
# 1) max absolute error / relative error
max_err = np.abs(ort_out - pt_out).max()
# 2) cosine similarity
cos_sim = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
# 3) business metrics (the hardcore check): AUC / top-1 acc before vs after conversionRules of thumb (commonly used in industry):
| Comparison | Max absolute error | Cosine similarity | Notes |
|---|---|---|---|
| Same framework FP32 vs FP32 | < 1e-4 | > 0.9999 | Pure operator implementation differences |
| PyTorch vs ONNX Runtime (FP32) | < 1e-3 | > 0.999 | Usually acceptable |
| FP32 vs INT8 quantized | < 1% metric drop vs baseline | Task-dependent | Quantization needs separate calibration, see Quantization |
The key discipline: validate on samples from the same distribution as production traffic—not hand-picked "nice-looking" training-set samples—and judge by business metrics, not bit-for-bit alignment.
7. GGUF and LLM Quantization
GGUF (defined by the llama.cpp community) plays a special role in LLM deployment:
- Single self-contained file: weights + tokenizer + model structure + hyperparameters all bundled—download and run;
- Embedded quantization schemes: preset formats like
q4_K_Mandq8_0(see Quantization) quantize weights block by block—no user-side calibration needed; - CPU as a first-class citizen: runs even without a GPU, the de facto standard for local LLMs (Ollama, LM Studio, etc.).
But it has two limitations: the quantization scheme is tightly bound to the inference runtime (GGUF's ggml tensor layout is optimized mainly for llama.cpp), and it's unfriendly to high-performance inference stacks like vLLM/TensorRT-LLM. Production LLM services therefore usually skip GGUF and use FP16/BF16 or dedicated AWQ/GPTQ quantization (see LLM Inference Optimization).
Trade-offs
| Decision point | Options | When to use |
|---|---|---|
| Portability vs peak performance | ONNX (portable) vs TensorRT (hardware-bound) | Multi-platform wants ONNX; single-machine peak performance wants TensorRT |
| Direct from framework vs intermediate format | Direct is easier; intermediate formats are portable | Cross-framework/cross-hardware requires an intermediate format |
| When to export | Rush before launch vs smoke test during training | Always the latter—see the MLOps gate design |
| GGUF vs dedicated quantization for LLMs | GGUF (llama.cpp ecosystem) vs AWQ/GPTQ (production serving) | GGUF for local tools; AWQ/GPTQ for high-concurrency services |
One-line summary: model formats are the "language" of deployment—decide the export path based on target hardware and scenario, and always validate after converting; an unvalidated export is as good as none.
Further Reading
- Quantization — the most common next optimization after export; INT8 requires re-export and calibration
- GPUs and Hardware Selection — formats ultimately land on a specific card's compute and bandwidth
- Inference: From Forward Pass to Inference Engines — how engines consume ONNX and optimize the graph
- TensorRT and Edge Deployment — a complete walkthrough from ONNX to .engine
- Common Pitfalls and Anti-Patterns — export-related failures seen in the field