Skip to content

Model Formats and Conversion

At a glance From PyTorch's .pt to ONNX, TensorRT, TFLite, and GGUF—the model format determines whether you can run efficient inference on the target hardware. This article surveys the mainstream formats, the conversion pipeline, operator compatibility pitfalls, and post-conversion numerical validation.

Model Formats and Conversion ​

The one-sentence definition: a model format is the serialization container for weights and the computation graph—to run efficient inference on the target hardware, the model first has to be "translated" into a language that hardware understands, and that translation is model conversion.

Industry insight: the model format is the most underrated link in the deployment chain. The .pt file PyTorch produces carries Python runtime dependencies; forget one dynamic_axes entry when exporting to ONNX and dynamic batching fails outright in production; forget to re-export after quantization and your INT8 speedup evaporates. Industry experience: a sizable share of deployment incidents aren't "the model is bad"—they're "wrong format, bad export, missing validation." Understanding the format ecosystem is laying the foundation for the entire deployment chain.

1. Why an Intermediate Format Is Needed ​

1.1 Three Decouplings ​

  • Decoupling from the training framework: PyTorch is great for iterating during training, but production wants no Python or framework dependencies—you can't drag a 2GB PyTorch environment around just to run one inference on a 500MB model.
  • Decoupling from hardware: the same model has to run on NVIDIA GPUs (TensorRT), Intel/AMD CPUs (OpenVINO/ONNX Runtime), phones (TFLite/Core ML), and browsers (WebAssembly)—you can't make every hardware target integrate with PyTorch directly.
  • Decoupling from the inference runtime: ONNX is a standard intermediate representation (IR) that ONNX Runtime, Triton, and TensorRT can all consume, avoiding vendor lock-in.

1.2 Two Conversion Routes ​

text
Route 1 (the standard route):
  PyTorch .pt ──export──► ONNX ──optimize/compile──► TensorRT .engine / OpenVINO IR
  (developer-friendly, portable)     (specialized to the target hardware, peak performance)

Route 2 (the direct route):
  PyTorch .pt ──torch.compile / direct load──► PyTorch runtime
  (skips the intermediate format, sacrifices some performance or portability)

Engineering practice favors the standard route: with ONNX as a stable intermediate layer, you can validate correctness quickly during development and then recompile for each target hardware. In the LLM era, GGUF is a special case of "the format is the deployment"—see Section 6 of this chapter.

2. The Mainstream Format Landscape ​

FormatEcosystemTarget hardwareOne-line summary
.pt / .pthPyTorchGeneral (needs PyTorch runtime)Training artifact, holds weights and structure (or a state_dict)
TorchScriptPyTorchGeneral (runs in C++ too)PyTorch's serialized graph format, no external dependencies
ONNXOpen standardCross-framework, cross-hardwareThe de facto interchange format, broadest ecosystem
.engineNVIDIA TensorRTNVIDIA GPUsCompiled artifact, peak performance, hardware-bound
.tfliteTensorFlow LiteMobile/edge/microcontrollersDe facto mobile standard, supports quantization
Core MLAppleApple siliconApple ecosystem's native format
OpenVINO IRIntelIntel CPU/GPU/NPUOptimized for Intel hardware, a strong CPU inference tool
GGUFllama.cpp ecosystemBoth CPU and GPULLM-specific, embeds quantization schemes and tokenizer
SafeTensorsHugging FaceGeneralFast, safe weight container, no code execution

Remember Three Watersheds

  • ONNX: the portability watershed—use it and your model isn't tied to any framework;
  • TensorRT .engine / OpenVINO IR: the performance watershed—use it and you're bound to one specific piece of hardware;
  • GGUF / TFLite: the scenario watershed—GGUF for running LLMs locally, TFLite for running on-device.

SafeTensors: Why the LLM Community Is Abandoning .bin ​

Hugging Face's pytorch_model.bin is essentially pickle, which executes arbitrary code on load—download a malicious weight file and you're poisoned. SafeTensors is a pure data format (JSON header + binary tensors): fast to load (zero deserialization overhead; loading 7B weights takes seconds) and safe. It's now the recommended container for LLM weight distribution. For the related supply-chain risks, see Security, Privacy, and Compliance.

3. The Standard Conversion Pipeline ​

text
Training output .pt
   │ ① export
   ▼
ONNX graph (with dynamic axes, opset version)
   │ ② optimize        graph optimization / precision calibration (INT8 needs calibration)
   ▼
Optimized ONNX
   │ ③ build           targeting the destination hardware
   ▼
TensorRT engine / OpenVINO IR / TFLite
   │ ④ validate        numerical comparison + measured performance
   ▼
Production

Every step can introduce problems—here they are one by one.

4. ONNX Export in Detail ​

4.1 Using torch.onnx.export Correctly ​

python
import torch
import torch.onnx

model = MyModel().eval().cuda()
dummy = torch.randn(1, 3, 224, 224).cuda()  # input shape must match production

torch.onnx.export(
    model,
    dummy,
    "model.onnx",
    input_names=["input"],      # name them; makes backend binding easier
    output_names=["logits"],
    dynamic_axes={              # critical: declare dynamic axes
        "input": {0: "batch", 2: "height", 3: "width"},
        "logits": {0: "batch"},
    },
    opset_version=17,           # opset version, see below
    do_constant_folding=True,   # constant folding, on by default
)

Three frequent mistakes:

  1. Forgetting eval(): BatchNorm/Dropout behave differently in training mode, so the exported graph is wrong.
  2. Missing dynamic_axes: shapes get frozen to dummy's by default, and batch=8 in production fails outright. Declare every axis that might vary.
  3. Export with realistic input shapes: shape is a hard constraint at export time; a mismatch with the real input leads to a broken graph.

4.2 What the Opset Version Is ​

The ONNX spec evolves operator semantics version by version; opset_version=17 means you're using the 17th operator set. Rule of thumb: use the highest version that covers all your operators and is supported by the target runtime. ONNX Runtime maintains an official operator compatibility matrix—for example, some new semantics from opset 18 aren't recognized by older runtime versions. The practical workflow:

text
Check the runtime's docs → pick the highest opset it supports → on "unsupported operator" errors, lower the version or rewrite the graph

4.3 Validate After Export with onnx.checker and onnxruntime ​

python
import onnx
import onnxruntime as ort

onnx.checker.check_model("model.onnx")          # structural validity
sess = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider"])
# feed a few real inputs and compare ONNX outputs vs PyTorch outputs (see Section 6 for numerical validation)

5. Operator Incompatibility: The Most Common Deployment Trap ​

5.1 What the Problem Looks Like ​

Export to ONNX fails, or building on the target engine reports "Unsupported operator / op type not registered." Common causes:

  • Dynamic control flow: if / for that depend on tensor values—ONNX doesn't support this; rewrite using torch.where or fixed loops.
  • Training-only operators: F.dropout (fixed by eval mode), anything torch.autograd-related.
  • Custom / new operators: e.g., older versions don't support torch.roll, and some attention implementations.
  • Unsupported dtypes: some engines reject FP64 or int64 inputs.

5.2 Three Mitigation Strategies (Cheapest First) ​

StrategyHowCostWhen to use
Replace the operatorRewrite with equivalent standard ops (e.g., attention as matmul+softmax)LowMost cases
Partition the graphRun unsupported subgraphs as Python fallback; everything else on the engineMediumA few stubborn operators
Fall back to CPUForce that operator to run on CPULowRare operators, ultra-low-traffic paths

The Cost of Partitioning and Fallbacks

Once the graph is split, intermediate tensors get copied back and forth, and performance can drop 20%+. Replace operators first; partition only as a last resort; reserve CPU fallback for paths that "run a handful of times a year."

5.3 The Real Fix: The Newer the Operators, the Earlier You Validate ​

New models (especially novel Transformer architectures) should get an export smoke test during training—not a week before launch, when you discover the engine doesn't support some attention operator. This belongs to the "model deployability check" gate in the MLOps Deployment Pipeline.

6. Post-Conversion Numerical Validation ​

6.1 Why Validation Is Non-Negotiable ​

Export/compilation doesn't change the math, but operator reordering, low-precision kernels, and different accumulation orders all introduce numerical differences. If you skip validation and ship, then when business metrics go wrong you can't tell whether it's the model or the conversion.

6.2 Three-Layer Validation ​

python
import numpy as np

# 1) max absolute error / relative error
max_err = np.abs(ort_out - pt_out).max()
# 2) cosine similarity
cos_sim = np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
# 3) business metrics (the hardcore check): AUC / top-1 acc before vs after conversion

Rules of thumb (commonly used in industry):

ComparisonMax absolute errorCosine similarityNotes
Same framework FP32 vs FP32< 1e-4> 0.9999Pure operator implementation differences
PyTorch vs ONNX Runtime (FP32)< 1e-3> 0.999Usually acceptable
FP32 vs INT8 quantized< 1% metric drop vs baselineTask-dependentQuantization needs separate calibration, see Quantization

The key discipline: validate on samples from the same distribution as production traffic—not hand-picked "nice-looking" training-set samples—and judge by business metrics, not bit-for-bit alignment.

7. GGUF and LLM Quantization ​

GGUF (defined by the llama.cpp community) plays a special role in LLM deployment:

  • Single self-contained file: weights + tokenizer + model structure + hyperparameters all bundled—download and run;
  • Embedded quantization schemes: preset formats like q4_K_M and q8_0 (see Quantization) quantize weights block by block—no user-side calibration needed;
  • CPU as a first-class citizen: runs even without a GPU, the de facto standard for local LLMs (Ollama, LM Studio, etc.).

But it has two limitations: the quantization scheme is tightly bound to the inference runtime (GGUF's ggml tensor layout is optimized mainly for llama.cpp), and it's unfriendly to high-performance inference stacks like vLLM/TensorRT-LLM. Production LLM services therefore usually skip GGUF and use FP16/BF16 or dedicated AWQ/GPTQ quantization (see LLM Inference Optimization).

Trade-offs ​

Decision pointOptionsWhen to use
Portability vs peak performanceONNX (portable) vs TensorRT (hardware-bound)Multi-platform wants ONNX; single-machine peak performance wants TensorRT
Direct from framework vs intermediate formatDirect is easier; intermediate formats are portableCross-framework/cross-hardware requires an intermediate format
When to exportRush before launch vs smoke test during trainingAlways the latter—see the MLOps gate design
GGUF vs dedicated quantization for LLMsGGUF (llama.cpp ecosystem) vs AWQ/GPTQ (production serving)GGUF for local tools; AWQ/GPTQ for high-concurrency services

One-line summary: model formats are the "language" of deployment—decide the export path based on target hardware and scenario, and always validate after converting; an unvalidated export is as good as none.

Further Reading ​

References ​