Appearance
ONNX Runtime: Cross-Platform
1. Definition: Freeing the "Model" from the "Framework"
ONNX Runtime (ORT) is the cross-platform inference engine that Microsoft has open-sourced and maintained since 2018. It takes a model file in ONNX (Open Neural Network Exchange) format as input and produces inference results that "run anywhere." Its core promise in one sentence: one model, run everywhere.
ONNX is an open intermediate representation (IR) jointly proposed by Facebook and Microsoft in 2017, aimed at breaking down framework barriers — a model trained in PyTorch can be exported to ONNX and run in C++, Java, C#, JS, on mobile, and on WebGPU. Understanding the relationship between ONNX and ORT is key:
┌──────────────────────────────────────────────────────────────┐
│ Training frameworks (PyTorch / TF / JAX / PaddlePaddle) │
└───────────────────────┬──────────────────────────────────────┘
│ export (torch.onnx.export / tf2onnx)
▼
┌──────────────────┐
│ ONNX IR (.onnx) │ ← standard operator set + standard format
└────────┬─────────┘
│ ONNX Runtime loads
▼
┌───────────────────────────────────────────────────┐
│ Execution Provider selection │
│ CPU / CUDA / TensorRT / DirectML / │
│ CoreML / QNN / OpenVINO / DML / NNAPI │
└──────────────────┬────────────────────────────────┘
│
▼
Inference result (a single run may span multiple EPs)The relationship between ORT and TensorRT deserves its own explanation: TensorRT is "the flagship on GPU," ORT is "the workhorse across all platforms." ORT has its own CUDA EP and can also dispatch nodes to the TensorRT EP for acceleration — it is the upper-layer scheduler, and TensorRT is one of its "GPU accelerators."
2. ONNX IR: A Standard Operator Set + Computation Graph
ONNX defines its IR with protobuf. The core trio:
- Node: one operator call, with op_type (e.g.,
Conv,MatMul), attributes (e.g.,kernel_size=3), and inputs/outputs. - Tensor: weights and activations, with dtype (FP32/FP16/INT8/...) and shape.
- Graph: a DAG of nodes + tensors + initializers (weights).
Every operator is defined in an ONNX Opset. The opset version number matters — a higher version supports more operators, but runtimes do not necessarily support them all. When exporting, the advice is "use the lowest opset that suffices" (PyTorch models usually work fine with opset 14–17).
python
# Validate an ONNX model
import onnx
model = onnx.load("model.onnx")
onnx.checker.check_model(model) # static check of IR validity
print(model.opset_import[0].version) # opset version
# Simplify the model (fold away redundant nodes produced by constant folding / shape inference)
import onnxsim
model_sim, ok = onnxsim.simplify(model)
onnx.save(model_sim, "model_sim.onnx")onnxsim is a must-have tool: it folds away statically determinable nodes such as Shape, Gather, and Reshape. The simplified model is usually easier for the runtime to optimize.
3. Execution Provider: The Pluggable Hardware Backend
ORT achieves "cross-platform" through the Execution Provider (EP) mechanism. Each EP wraps a set of hardware-implemented operators, and when loading a model, ORT assigns each node to the most suitable EP based on capability sets.
| EP | Platform / Hardware | Typical use case |
|---|---|---|
| CPU EP | Default on all platforms | Universal fallback; AVX2/AVX512 enabled automatically |
| CUDA EP | NVIDIA GPU | 5–20× faster than CPU EP |
| TensorRT EP | NVIDIA GPU | Hands fusable subgraphs to TRT on top of CUDA EP (see TensorRT) |
| DirectML EP | Windows + any DX12 GPU | Cross-vendor Windows inference (AMD/Intel/NVIDIA) |
| CoreML EP | iOS / macOS | Apple Neural Engine acceleration |
| NNAPI EP | Android | Android NPU/DSP |
| QNN EP | Qualcomm Hexagon | High-performance path to Snapdragon NPUs |
| OpenVINO EP | Intel CPU/iGPU | See OpenVINO |
| ROCm EP | AMD GPU | ROCm stack |
| XNNPACK / WebGPU | Mobile / Web | Lightweight + browser |
Node Assignment and Fallback
EP assignment is "greedy + capability matching." The default order is: developer-declared priority → EP capability. Nodes not claimed by any accelerated EP fall back to the CPU EP — this is ORT's compatibility safety net, but the price is a performance hit from "context switches" across EPs.
python
import onnxruntime as ort
# Set EP priority
session = ort.InferenceSession(
"model.onnx",
providers=[
("CUDAExecutionProvider", {
"device_id": 0,
"arena_extend_strategy": "kSameAsRequested",
"gpu_mem_limit": 8 * 1024 * 1024 * 1024, # 8GB
}),
"TensorrtExecutionProvider", # second priority
"CPUExecutionProvider", # fallback
],
)
# Check the actual assignment
print(session.get_providers()) # actually enabled EPs
print(session._get_provider_options()) # config of each EPEP priority: more aggressive is not always better
Aggressively enabling the TensorRT EP lets TRT grab some nodes and run them inside TRT subgraphs; data then has to be copied (memcpy) between the subgraphs and the CPU EP, which can actually slow things down. First inspect the subgraph partition with the trt_plugin_creation tool or onnxruntime perf_analyzer, then decide which nodes to force onto CPU/CUDA.
4. The Graph Optimization Pipeline (GRPO)
Internally, ORT runs GRPO (Graph Optimizer Pipeline) for graph optimization, divided into several levels:
- Basic level (L1–L3): constant folding, shape inference, redundant operator elimination, Cast elimination. No hardware dependencies; all EPs benefit.
- Extended level (L4–L6): fused-operator replacement (Conv+BN+Relu → FusedConv), subgraph replacement. Depends on EP capabilities.
- Layout optimizer: NCHW ↔ NHWC conversion, choosing layouts per EP preference.
- Quantization / Pre-processor: works with the quantization tool
onnxruntime.quantization.
See Computation Graph Optimization and Kernel Fusion and Custom Kernels. A typical fusion:
Original: Conv → Add(bias) → Relu
Optimized: FusedConv (Conv+Add+Relu in one kernel call)5. Quantization: Built-in Dynamic / Static Quantization
ORT ships with quantization APIs; see Model Quantization Fundamentals.
python
from onnxruntime.quantization import quantize_dynamic, quantize_static, QuantType, QuantFormat
# Dynamic quantization (INT8 weights, activations quantized at runtime, zero-shot)
quantize_dynamic(
model_input="model.onnx",
model_output="model_int8_dynamic.onnx",
weight_type=QuantType.QInt8,
)
# Static quantization (INT8 weights + activations, requires calibration data)
quantize_static(
model_input="model.onnx",
model_output="model_int8_static.onnx",
calibration_data_reader=MyDataReader(...),
quant_format=QuantFormat.QDQ, # QDQ has the best portability
per_channel=True,
)- Dynamic quantization: zero calibration data; only weights are quantized, and activation scales are computed at inference time. CPU-friendly, little gain on GPU.
- Static quantization: requires calibration data; quantizes weights + activations, with better performance, but accuracy loss must be evaluated.
- QDQ (QuantizeLinear/DequantizeLinear) vs QOperator: QDQ expresses quantization with standard ONNX operators and has the best cross-EP compatibility; QOperator uses dedicated quantized operators — slightly faster but less compatible.
6. The Cross-Framework Toolchain
Turning a PyTorch model into ONNX is the most common pipeline:
python
import torch
class Model(torch.nn.Module):
def forward(self, x):
return torch.relu(self.fc(x))
model = Model().eval()
dummy = torch.randn(1, 3, 224, 224)
torch.onnx.export(
model, (dummy,),
"model.onnx",
input_names=["input"],
output_names=["output"],
opset_version=17,
dynamic_axes={"input": {0: "batch"}, "output": {0: "batch"}}, # dynamic batch
dynamo=True, # recommended for PyTorch 2.x: Dynamo-based export, more robust
)
# Validate
import onnxruntime as ort
import numpy as np
sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])
np.testing.assert_allclose(
sess.run(None, {"input": dummy.numpy()})[0],
model(dummy).detach().numpy(),
rtol=1e-3, atol=1e-4,
)Common export pitfalls
- Too many dynamic dimensions: more dynamic_axes is not better. Fully dynamic shapes prevent ORT from pre-allocating its arena and hurt performance; make only batch / seq_len dynamic.
- Opset too new: operators from opset 18+ are not recognized by older ORT. Lock opset 14–17 + ORT 1.16+ for production.
- Unsupported operators:
torch.nn.functional.ctc_lossand custom LSTM variants are often unsupported — either switch implementations or write an ONNX Custom Op. dynamo=Truevs the legacytorch.onnx.export: PyTorch 2.9+ recommends dynamo, but the legacy path is more robust for certain dynamic control-flow scenarios; the two will coexist for a while.
The full toolchain:
- torch.onnx.export: PyTorch → ONNX
- tf2onnx: TF → ONNX
- onnxsim: model simplification
- onnx.checker + onnx.shape_inference: static checks
- onnxruntime.tools.onnx_test_runner: cross-platform testing
- onnxruntime.transformers: BERT/Transformer optimization tools
- netron: visualize ONNX model structure
7. Performance Data: The "Counterintuitive" Part
ORT's counterintuitive advantage is on CPU: in many scenarios, the ONNX Runtime CPU EP is even faster than native PyTorch on CPU — because ORT uses MKL-DNN (oneDNN), AVX512, and fused GEMM, while PyTorch's default CPU path is less optimized.
| Model | Hardware | Precision | PyTorch CPU | ONNX Runtime CPU | Speedup |
|---|---|---|---|---|---|
| BERT-base | Xeon 8358 | FP32 | 18 ms | 11 ms | 1.6× |
| ResNet-50 | Xeon 8358 | INT8 | 8 ms | 4.5 ms | 1.8× |
| MobileNetV2 | Xeon 8358 | FP32 | 4.0 ms | 2.2 ms | 1.8× |
| YOLOv8-m | Xeon 8358 | FP16 | 22 ms | 15 ms | 1.5× |
On GPU, ORT's CUDA EP trails TensorRT by 1.2–1.8× (TensorRT goes deeper in fusion and kernel tuning), but the development barrier is much lower. See Inference Engine Comparison.
8. Limitations and Boundaries
- Lagging support for new operators: custom attention and new activation functions lag behind in the ONNX opset; you have to wait for an opset upgrade or write a Custom Op.
- Falls short of TensorRT for peak GPU performance: on NVIDIA GPUs, if you can run the TensorRT EP, don't settle for the CUDA EP alone — but the TensorRT EP also costs more build time and memory.
- Weak for large models (LLMs): ORT itself lacks PagedAttention, continuous batching, and other LLM-critical optimizations — for LLM inference, use vLLM or TensorRT-LLM. ORT's strength is small-to-medium models and cross-platform deployment.
- Limited dynamic control flow: ONNX's native
If/Loopsupport is limited; data-dependent branches like PyTorch'sif x.sum() > 0often fail to export. - Switching overhead between EPs: subgraph boundaries across EPs incur memcpy; node assignment policies must be tuned with care.
9. Comparison with Similar Solutions
| Solution | Relationship with ORT |
|---|---|
| TensorRT | ORT's TensorRT EP calls TRT to accelerate subgraphs; ORT is the upper layer, TRT the backend |
| OpenVINO | The OpenVINO EP lets ORT run on Intel hardware; you can also use the OpenVINO Runtime standalone (deeper optimization for Intel CPUs) |
| Triton Inference Server | Triton's onnxruntime backend is literally ORT; Triton is the serving layer, ORT the engine |
| llama.cpp | Both lean "cross-platform," but llama.cpp is for LLMs, ORT for small-to-medium models |
| Mobile Deployment | ORT Mobile is the trimmed-down mobile version of ORT |
10. Where to Go Next
- Concept pages: Computation Graph Optimization, Kernel Fusion and Custom Kernels, Model Quantization Fundamentals, GPU Architecture and Optimization, Latency, Throughput, and Concurrency
- Case-study pages: TensorRT, OpenVINO, Triton Inference Server, Mobile Deployment
- Practice pages: Inference Engine Comparison, Tuning and Performance Optimization, Inference Benchmarking in Practice, Common Pitfalls and Anti-Patterns
- Resource pages: Glossary, Hardware Primer, Curated Resources
References
- ONNX Runtime official documentation — the authority on performance and APIs
- ONNX IR Spec — IR definition
- ONNX Operators — operator specification
- Microsoft. ONNX Runtime Performance Tuning — tuning guide
- onnxsim — model simplification tool
- tf2onnx — TF to ONNX
- Netron — online ONNX visualization