Skip to content

ONNX Runtime: Cross-Platform

At a glance ONNX Runtime takes "model as a file" to the extreme — a single .onnx file runs on CPU, CUDA, DirectML, CoreML, and QNN, powered by the Execution Provider mechanism and a rigorous graph optimization pipeline. This article dissects the ONNX IR, the EP mechanism, the GRPO optimizer, and the cross-framework toolchain.

ONNX Runtime: Cross-Platform ​

1. Definition: Freeing the "Model" from the "Framework" ​

ONNX Runtime (ORT) is the cross-platform inference engine that Microsoft has open-sourced and maintained since 2018. It takes a model file in ONNX (Open Neural Network Exchange) format as input and produces inference results that "run anywhere." Its core promise in one sentence: one model, run everywhere.

ONNX is an open intermediate representation (IR) jointly proposed by Facebook and Microsoft in 2017, aimed at breaking down framework barriers — a model trained in PyTorch can be exported to ONNX and run in C++, Java, C#, JS, on mobile, and on WebGPU. Understanding the relationship between ONNX and ORT is key:

┌──────────────────────────────────────────────────────────────┐
│  Training frameworks (PyTorch / TF / JAX / PaddlePaddle)     │
└───────────────────────┬──────────────────────────────────────┘
                        │ export (torch.onnx.export / tf2onnx)
                        ▼
               ┌──────────────────┐
               │  ONNX IR (.onnx) │  ← standard operator set + standard format
               └────────┬─────────┘
                        │ ONNX Runtime loads
                        ▼
   ┌───────────────────────────────────────────────────┐
   │          Execution Provider selection              │
   │  CPU / CUDA / TensorRT / DirectML /                │
   │  CoreML / QNN / OpenVINO / DML / NNAPI             │
   └──────────────────┬────────────────────────────────┘
                      │
                      ▼
        Inference result (a single run may span multiple EPs)

The relationship between ORT and TensorRT deserves its own explanation: TensorRT is "the flagship on GPU," ORT is "the workhorse across all platforms." ORT has its own CUDA EP and can also dispatch nodes to the TensorRT EP for acceleration — it is the upper-layer scheduler, and TensorRT is one of its "GPU accelerators."

2. ONNX IR: A Standard Operator Set + Computation Graph ​

ONNX defines its IR with protobuf. The core trio:

  • Node: one operator call, with op_type (e.g., Conv, MatMul), attributes (e.g., kernel_size=3), and inputs/outputs.
  • Tensor: weights and activations, with dtype (FP32/FP16/INT8/...) and shape.
  • Graph: a DAG of nodes + tensors + initializers (weights).

Every operator is defined in an ONNX Opset. The opset version number matters — a higher version supports more operators, but runtimes do not necessarily support them all. When exporting, the advice is "use the lowest opset that suffices" (PyTorch models usually work fine with opset 14–17).

python
# Validate an ONNX model
import onnx

model = onnx.load("model.onnx")
onnx.checker.check_model(model)              # static check of IR validity
print(model.opset_import[0].version)         # opset version

# Simplify the model (fold away redundant nodes produced by constant folding / shape inference)
import onnxsim
model_sim, ok = onnxsim.simplify(model)
onnx.save(model_sim, "model_sim.onnx")

onnxsim is a must-have tool: it folds away statically determinable nodes such as Shape, Gather, and Reshape. The simplified model is usually easier for the runtime to optimize.

3. Execution Provider: The Pluggable Hardware Backend ​

ORT achieves "cross-platform" through the Execution Provider (EP) mechanism. Each EP wraps a set of hardware-implemented operators, and when loading a model, ORT assigns each node to the most suitable EP based on capability sets.

EPPlatform / HardwareTypical use case
CPU EPDefault on all platformsUniversal fallback; AVX2/AVX512 enabled automatically
CUDA EPNVIDIA GPU5–20× faster than CPU EP
TensorRT EPNVIDIA GPUHands fusable subgraphs to TRT on top of CUDA EP (see TensorRT)
DirectML EPWindows + any DX12 GPUCross-vendor Windows inference (AMD/Intel/NVIDIA)
CoreML EPiOS / macOSApple Neural Engine acceleration
NNAPI EPAndroidAndroid NPU/DSP
QNN EPQualcomm HexagonHigh-performance path to Snapdragon NPUs
OpenVINO EPIntel CPU/iGPUSee OpenVINO
ROCm EPAMD GPUROCm stack
XNNPACK / WebGPUMobile / WebLightweight + browser

Node Assignment and Fallback ​

EP assignment is "greedy + capability matching." The default order is: developer-declared priority → EP capability. Nodes not claimed by any accelerated EP fall back to the CPU EP — this is ORT's compatibility safety net, but the price is a performance hit from "context switches" across EPs.

python
import onnxruntime as ort

# Set EP priority
session = ort.InferenceSession(
    "model.onnx",
    providers=[
        ("CUDAExecutionProvider", {
            "device_id": 0,
            "arena_extend_strategy": "kSameAsRequested",
            "gpu_mem_limit": 8 * 1024 * 1024 * 1024,  # 8GB
        }),
        "TensorrtExecutionProvider",   # second priority
        "CPUExecutionProvider",        # fallback
    ],
)

# Check the actual assignment
print(session.get_providers())                # actually enabled EPs
print(session._get_provider_options())        # config of each EP

EP priority: more aggressive is not always better

Aggressively enabling the TensorRT EP lets TRT grab some nodes and run them inside TRT subgraphs; data then has to be copied (memcpy) between the subgraphs and the CPU EP, which can actually slow things down. First inspect the subgraph partition with the trt_plugin_creation tool or onnxruntime perf_analyzer, then decide which nodes to force onto CPU/CUDA.

4. The Graph Optimization Pipeline (GRPO) ​

Internally, ORT runs GRPO (Graph Optimizer Pipeline) for graph optimization, divided into several levels:

  • Basic level (L1–L3): constant folding, shape inference, redundant operator elimination, Cast elimination. No hardware dependencies; all EPs benefit.
  • Extended level (L4–L6): fused-operator replacement (Conv+BN+Relu → FusedConv), subgraph replacement. Depends on EP capabilities.
  • Layout optimizer: NCHW ↔ NHWC conversion, choosing layouts per EP preference.
  • Quantization / Pre-processor: works with the quantization tool onnxruntime.quantization.

See Computation Graph Optimization and Kernel Fusion and Custom Kernels. A typical fusion:

Original:   Conv → Add(bias) → Relu
Optimized:  FusedConv (Conv+Add+Relu in one kernel call)

5. Quantization: Built-in Dynamic / Static Quantization ​

ORT ships with quantization APIs; see Model Quantization Fundamentals.

python
from onnxruntime.quantization import quantize_dynamic, quantize_static, QuantType, QuantFormat

# Dynamic quantization (INT8 weights, activations quantized at runtime, zero-shot)
quantize_dynamic(
    model_input="model.onnx",
    model_output="model_int8_dynamic.onnx",
    weight_type=QuantType.QInt8,
)

# Static quantization (INT8 weights + activations, requires calibration data)
quantize_static(
    model_input="model.onnx",
    model_output="model_int8_static.onnx",
    calibration_data_reader=MyDataReader(...),
    quant_format=QuantFormat.QDQ,   # QDQ has the best portability
    per_channel=True,
)
  • Dynamic quantization: zero calibration data; only weights are quantized, and activation scales are computed at inference time. CPU-friendly, little gain on GPU.
  • Static quantization: requires calibration data; quantizes weights + activations, with better performance, but accuracy loss must be evaluated.
  • QDQ (QuantizeLinear/DequantizeLinear) vs QOperator: QDQ expresses quantization with standard ONNX operators and has the best cross-EP compatibility; QOperator uses dedicated quantized operators — slightly faster but less compatible.

6. The Cross-Framework Toolchain ​

Turning a PyTorch model into ONNX is the most common pipeline:

python
import torch

class Model(torch.nn.Module):
    def forward(self, x):
        return torch.relu(self.fc(x))

model = Model().eval()
dummy = torch.randn(1, 3, 224, 224)

torch.onnx.export(
    model, (dummy,),
    "model.onnx",
    input_names=["input"],
    output_names=["output"],
    opset_version=17,
    dynamic_axes={"input": {0: "batch"}, "output": {0: "batch"}},  # dynamic batch
    dynamo=True,  # recommended for PyTorch 2.x: Dynamo-based export, more robust
)

# Validate
import onnxruntime as ort
import numpy as np

sess = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])
np.testing.assert_allclose(
    sess.run(None, {"input": dummy.numpy()})[0],
    model(dummy).detach().numpy(),
    rtol=1e-3, atol=1e-4,
)

Common export pitfalls

  1. Too many dynamic dimensions: more dynamic_axes is not better. Fully dynamic shapes prevent ORT from pre-allocating its arena and hurt performance; make only batch / seq_len dynamic.
  2. Opset too new: operators from opset 18+ are not recognized by older ORT. Lock opset 14–17 + ORT 1.16+ for production.
  3. Unsupported operators: torch.nn.functional.ctc_loss and custom LSTM variants are often unsupported — either switch implementations or write an ONNX Custom Op.
  4. dynamo=True vs the legacy torch.onnx.export: PyTorch 2.9+ recommends dynamo, but the legacy path is more robust for certain dynamic control-flow scenarios; the two will coexist for a while.

The full toolchain:

  • torch.onnx.export: PyTorch → ONNX
  • tf2onnx: TF → ONNX
  • onnxsim: model simplification
  • onnx.checker + onnx.shape_inference: static checks
  • onnxruntime.tools.onnx_test_runner: cross-platform testing
  • onnxruntime.transformers: BERT/Transformer optimization tools
  • netron: visualize ONNX model structure

7. Performance Data: The "Counterintuitive" Part ​

ORT's counterintuitive advantage is on CPU: in many scenarios, the ONNX Runtime CPU EP is even faster than native PyTorch on CPU — because ORT uses MKL-DNN (oneDNN), AVX512, and fused GEMM, while PyTorch's default CPU path is less optimized.

ModelHardwarePrecisionPyTorch CPUONNX Runtime CPUSpeedup
BERT-baseXeon 8358FP3218 ms11 ms1.6×
ResNet-50Xeon 8358INT88 ms4.5 ms1.8×
MobileNetV2Xeon 8358FP324.0 ms2.2 ms1.8×
YOLOv8-mXeon 8358FP1622 ms15 ms1.5×

On GPU, ORT's CUDA EP trails TensorRT by 1.2–1.8× (TensorRT goes deeper in fusion and kernel tuning), but the development barrier is much lower. See Inference Engine Comparison.

8. Limitations and Boundaries ​

  1. Lagging support for new operators: custom attention and new activation functions lag behind in the ONNX opset; you have to wait for an opset upgrade or write a Custom Op.
  2. Falls short of TensorRT for peak GPU performance: on NVIDIA GPUs, if you can run the TensorRT EP, don't settle for the CUDA EP alone — but the TensorRT EP also costs more build time and memory.
  3. Weak for large models (LLMs): ORT itself lacks PagedAttention, continuous batching, and other LLM-critical optimizations — for LLM inference, use vLLM or TensorRT-LLM. ORT's strength is small-to-medium models and cross-platform deployment.
  4. Limited dynamic control flow: ONNX's native If / Loop support is limited; data-dependent branches like PyTorch's if x.sum() > 0 often fail to export.
  5. Switching overhead between EPs: subgraph boundaries across EPs incur memcpy; node assignment policies must be tuned with care.

9. Comparison with Similar Solutions ​

SolutionRelationship with ORT
TensorRTORT's TensorRT EP calls TRT to accelerate subgraphs; ORT is the upper layer, TRT the backend
OpenVINOThe OpenVINO EP lets ORT run on Intel hardware; you can also use the OpenVINO Runtime standalone (deeper optimization for Intel CPUs)
Triton Inference ServerTriton's onnxruntime backend is literally ORT; Triton is the serving layer, ORT the engine
llama.cppBoth lean "cross-platform," but llama.cpp is for LLMs, ORT for small-to-medium models
Mobile DeploymentORT Mobile is the trimmed-down mobile version of ORT

10. Where to Go Next ​

References ​