Appearance
TensorRT and Edge Deployment: From ONNX to Jetson
One-line definition: TensorRT is NVIDIA's GPU inference engine. It compiles a trained model (ONNX, etc.) into an engine highly optimized for a specific GPU — layer fusion, precision calibration, and automatic kernel selection squeeze another order of magnitude out of inference latency — ready to deploy on edge devices like Jetson.
Why it's worth doing: GPUs on edge devices are an order of magnitude smaller than their cloud counterparts, so models have to be squeezed to the limit to run in real time. A Jetson Orin Nano has only 8–16GB of unified memory and a 7–25W power budget; native PyTorch may take 50ms+ to run ResNet50 on it, while TensorRT FP16 gets under 5ms and INT8 is faster still. This guide walks the full chain from ONNX preparation and trtexec conversion through INT8 calibration to Jetson deployment, with measured accuracy-versus-performance trade-offs.
1. What TensorRT Does
TensorRT does four things for GPU inference:
- Layer fusion (kernel fusion): fuses operators such as Conv+BN+ReLU into a single kernel, cutting kernel launches and memory traffic;
- Precision calibration: FP16/INT8 inference reduces bandwidth and compute (INT8 computes 2–4× faster and halves memory);
- Graph optimization and kernel selection: automatically picks the best kernels for the target GPU architecture (e.g. tensor-core paths for Ada/Lovelace);
- Memory optimization: static memory pools and fewer intermediate tensor allocations.
Verdict up front: TensorRT's payoff is inference latency — typically 2–8× faster than PyTorch eager. The price is that the compiled artifact is bound to a specific GPU model and precision, and INT8 requires calibration data. The underlying mechanics map to the "inference engines" section of Inference: from forward pass to inference engine.
2. The Optimization Pipeline at a Glance
text
PyTorch model
│ torch.onnx.export
▼
model.onnx
│ trtexec --fp16 / --int8 (+ calibration)
▼
model.engine (bound to GPU model and TensorRT version)
│ Triton backend / C++ Runtime / Python Runtime
▼
Runs on a Jetson or data-center GPU3. Step 1: Export and Validate the ONNX
python
import torch
import onnx
from torchvision.models import resnet50, ResNet50_Weights
model = resnet50(weights=ResNet50_Weights.DEFAULT).eval()
dummy = torch.randn(1, 3, 224, 224)
torch.onnx.export(
model, dummy, "resnet50.onnx",
input_names=["input"], output_names=["output"],
# Declare both dynamic batch and dynamic resolution explicitly (dynamic resolution is common on edge)
dynamic_axes={"input": {0: "batch", 2: "height", 3: "width"},
"output": {0: "batch", 2: "height", 3: "width"}},
opset_version=17,
)
# Validate the model structure with the onnx library
onnx.checker.check_model(onnx.load("resnet50.onnx"))
print("onnx ok, ops:", len(onnx.load("resnet50.onnx").graph.node))Export essentials and compatibility checks are covered in model formats and conversion. When export fails or accuracy shifts, suspect custom operators and the opset first.
4. Step 2: Conversion with trtexec (FP16 / INT8 / Dynamic Shapes)
trtexec is TensorRT's built-in conversion and benchmarking tool (it ships with TensorRT):
bash
# FP16: the usual choice, zero extra cost, roughly 2x the speed
trtexec --onnx=resnet50.onnx --saveEngine=resnet50_fp16.engine \
--fp16 \
--minShapes=input:1x3x224x224 \
--optShapes=input:1x3x224x224 \
--maxShapes=input:16x3x224x224
# INT8: requires a calibration set (see Section 5), roughly 3-4x the speed
trtexec --onnx=resnet50.onnx --saveEngine=resnet50_int8.engine \
--int8 \
--calib=/data/calib/calib.txt \
--calibCache=calib.cache \
--minShapes=input:1x3x224x224 \
--optShapes=input:1x3x224x224 \
--maxShapes=input:16x3x224x224The three dynamic-shape parameters, verdict first:
--minShapes/--optShapes/--maxShapesdefine the batch/resolution range:minsets the floor for memory preallocation,optis the tuning target for kernel selection (the shape you'll hit most often), andmaxcaps memory. The wider the range, the slower the build and the more memory reserved — keep it close to real usage;--calibCache: caches calibration results to a file so the second build doesn't redo calibration (otherwise every--int8build is slow);- The
.engineartifact is bound to a specific GPU model and TensorRT version: an engine built on an A10 won't run on a Jetson; changing hardware means rebuilding.
5. INT8 Calibration and Accuracy Validation
INT8 quantization needs a representative calibration set so TensorRT can measure per-layer activation distributions and derive scale factors. Straight to the point:
- The calibration set must represent real production input — use 500–2000 real images sampled from production, not a random slice of the training set (its distribution may have drifted from production);
- Always validate accuracy after calibration: run a validation set and compare metrics against FP32/FP16 (Top-1 or your task metric). If INT8 drops more than your threshold (e.g. >0.5% for image classification), switch to per-layer INT8+FP16 mixing (
--layer-precision) and only quantize the insensitive layers; - Accuracy vs. speed table (ResNet50, Jetson Orin Nano):
| Precision | Latency (ms) | Relative speed | Top-1 delta |
|---|---|---|---|
| FP32 | ~20 | 1.0× | baseline |
| FP16 | ~8 | ~2.5× | lossless |
| INT8 | ~4.5 | ~4.4× | approx. -0.2% to -0.5% |
For the general quantization trade-off (not just TensorRT), see quantization.
Per-Layer Precision Control: Plan B When INT8 Hurts Too Much
When whole-network INT8 degrades accuracy unacceptably, fall back to per-layer mixed precision: use --layer-precision to keep specific layers in FP16. Example command:
bash
trtexec --onnx=resnet50.onnx --saveEngine=resnet50_mixed.engine \
--int8 --fp16 \
--calib=/data/calib/calib.txt \
# Keep sensitive layers (e.g. the final fully-connected and normalization layers) in FP16
--layer-precision=fc:PFP16 \
--layer-precision=conv_5: PFP16Verdict first: try whole-network INT8 and measure the drop; if it's small, ship it; if it's large, tune per layer; if that still fails, consider QAT (quantization-aware training). Most CNNs lose <0.5% with whole-network INT8; Transformers and detection models concentrate sensitivity in specific layers, so per-layer control pays off visibly. You can dump the layer name list first with trtexec --dumpLayerInfo.
6. Runtime Inference: Loading the Engine with Python / C++
python
# trt_runtime.py
import tensorrt as trt
import numpy as np
from cuda import cudart # or pycuda
logger = trt.Logger(trt.Logger.WARNING)
runtime = trt.Runtime(logger)
# Deserialize the engine (binary; bound to the GPU it was built on)
with open("resnet50_fp16.engine", "rb") as f:
engine = runtime.deserialize_cuda_engine(f.read())
context = engine.create_execution_context()
# Dynamic shapes: the input shape must be set explicitly before each inference
context.set_input_shape("input", (1, 3, 224, 224))
# Allocate input/output device buffers (flow only; real code must handle streams and synchronization)
input_buf, output_buf, stream = allocate_buffers(engine, context)
# Copy input in -> context.execute_async_v2(bindings, stream) -> copy output outKey point: engines built with dynamic shapes require set_input_shape and buffer reallocation on every inference; in production you'll usually let Triton's TensorRT backend host the engine and skip hand-written memory management — it integrates seamlessly with Triton. The C++ path works the same way via InferRuntime; Jetson edge deployments commonly use C++, or Python + TensorRT from the Jetson-specific images.
Hosting the Engine with Triton: No Hand-Written Memory Management
When you run multiple engines in the data center or on a Jetson, Triton's tensorrt backend is the most trouble-free way to host them; just switch the backend to TensorRT in config.pbtxt:
protobuf
# models/resnet50/config.pbtxt
name: "resnet50"
backend: "tensorrt"
max_batch_size: 16
input [ { name: "input", data_type: TYPE_FP32, dims: [3, 224, 224] } ]
output [ { name: "output", data_type: TYPE_FP32, dims: [1000] } ]
instance_group [ { kind: KIND_GPU, count: 1 } ]Benefits: dynamic batching, concurrent scheduling, and metrics output all come for free, and engine file management is folded into the model repository. It integrates seamlessly with NVIDIA Triton multi-model serving; Triton parameters such as --backend-config=tensorrt,coalesce-request-input=true let you tune batching further.
7. Edge Deployment: Jetson
Jetson is NVIDIA's edge computing platform. A comparison of common models (note the unified memory):
| Model | Compute | Unified memory | Power | Typical use |
|---|---|---|---|---|
| Jetson Orin Nano | ~40 TOPS (INT8) | 8/16GB | 7–25W | Real-time detection at the camera |
| Jetson Orin NX | ~100 TOPS (INT8) | 16GB | 10–40W | Multi-stream video at the edge |
| Jetson AGX Orin | ~275 TOPS (INT8) | 64GB | 15–60W | Edge LLMs / heavy models |
Three deployment essentials:
- Use Jetson-specific images:
nvcr.io/nvidia/l4t-pytorch,l4t-tensorrt, or the JetPack SDK (which bundles a matching TensorRT version) — do not install data-center TensorRT on a Jetson; the architecture doesn't match; - Cross-compilation vs. on-device builds: the engine must be rebuilt on the target Jetson (or built offline with
trtexecon the board), and container images must be built foraarch64; - Power and thermal constraints: Jetson has power modes (
nvpmodel -m 0for max performance / low-power modes); pick the right mode for heavy models and monitor temperatures — thermal throttling causes latency spikes. For hardware selection and power trade-offs, see GPU and hardware selection.
The Real Challenges of Edge Deployment
Edge is not the cloud: the network can drop, devices can reboot, and nobody is standing by to restart the process. Engine files must be persisted locally (avoid recompiling at every boot — a single build can take minutes); the service needs to auto-start on boot and be restarted after crashes; and model updates should arrive via OTA rather than on-site reflashing. These "operational costs" are easy to overlook, yet they dominate edge work.
8. Measured Performance Comparison
Same Jetson Orin Nano, ResNet50, batch size 1, dynamic input 224×224:
| Setup | Latency (ms) | Throughput (FPS) | Memory usage |
|---|---|---|---|
| PyTorch eager (FP32) | ~22 | ~45 | ~500MB |
| TensorRT FP16 | ~8 | ~125 | ~300MB |
| TensorRT INT8 | ~4.5 | ~220 | ~250MB |
Conclusion: FP16 is the best-value starting point; move to INT8 once accuracy is acceptable. For mobile (no NVIDIA GPU), the alternative route is TFLite + quantization via TFLite/ML Kit; the strategy is the same (calibration + accuracy validation).
Common Pitfalls and Troubleshooting
| Pitfall | Symptom | Fix |
|---|---|---|
| Engine fails on a different GPU | engine file is generated on an incompatible device | Engines are bound to GPU/version; rebuild on the target device |
| Dynamic shape not set | Inference throws invalid shapes | Call set_input_shape before every inference |
| INT8 accuracy collapse | Top-1 drops 5%+ | Calibration set not representative; switch to per-layer mixed precision |
| Conversion fails | trtexec reports unsupported operators | Move to a higher opset, replace custom operators, add --onnx-normalization |
| Device memory OOM | Dynamic shape range set too wide | Narrow min/maxShapes; don't pad the max |
| Jetson can't keep up | Latency far above expectations | Check the nvpmodel power mode and thermal throttling |
| Version mismatch | Runtime reports missing trt version API | Use the same TensorRT version for the engine and the runtime |
Further Reading
- Inference: from forward pass to inference engine — the mechanics of TensorRT layer fusion and kernel selection
- Model formats and conversion — a full checklist for ONNX export, opsets, and operator compatibility
- Quantization — the theory behind INT8 calibration, PTQ/QAT, and accuracy trade-offs
- NVIDIA Triton multi-model serving — host engines with Triton's TensorRT backend and skip hand-written memory management
- GPU and hardware selection — the Jetson lineup, power modes, and edge hardware decisions
- Common pitfalls and antipatterns — edge operations pitfalls (OTA, auto-start on boot, crash restarts)
References
- NVIDIA TensorRT documentation: https://docs.nvidia.com/deeplearning/tensorrt/
- trtexec usage guide: https://docs.nvidia.com/deeplearning/tensorrt/developer-guide/index.html#trtexec
- NVIDIA Jetson modules: https://developer.nvidia.com/embedded/jetson-modules
- NVIDIA L4T / JetPack: https://developer.nvidia.com/embedded/jetpack