Skip to content

TensorRT and Edge Deployment

At a glance A hands-on path from optimizing a model into a TensorRT engine to deploying it on edge devices (Jetson): ONNX→TRT conversion, FP16/INT8 accuracy trade-offs, dynamic shapes, and edge deployment under power constraints.

TensorRT and Edge Deployment: From ONNX to Jetson ​

One-line definition: TensorRT is NVIDIA's GPU inference engine. It compiles a trained model (ONNX, etc.) into an engine highly optimized for a specific GPU — layer fusion, precision calibration, and automatic kernel selection squeeze another order of magnitude out of inference latency — ready to deploy on edge devices like Jetson.

Why it's worth doing: GPUs on edge devices are an order of magnitude smaller than their cloud counterparts, so models have to be squeezed to the limit to run in real time. A Jetson Orin Nano has only 8–16GB of unified memory and a 7–25W power budget; native PyTorch may take 50ms+ to run ResNet50 on it, while TensorRT FP16 gets under 5ms and INT8 is faster still. This guide walks the full chain from ONNX preparation and trtexec conversion through INT8 calibration to Jetson deployment, with measured accuracy-versus-performance trade-offs.

1. What TensorRT Does ​

TensorRT does four things for GPU inference:

  1. Layer fusion (kernel fusion): fuses operators such as Conv+BN+ReLU into a single kernel, cutting kernel launches and memory traffic;
  2. Precision calibration: FP16/INT8 inference reduces bandwidth and compute (INT8 computes 2–4× faster and halves memory);
  3. Graph optimization and kernel selection: automatically picks the best kernels for the target GPU architecture (e.g. tensor-core paths for Ada/Lovelace);
  4. Memory optimization: static memory pools and fewer intermediate tensor allocations.

Verdict up front: TensorRT's payoff is inference latency — typically 2–8× faster than PyTorch eager. The price is that the compiled artifact is bound to a specific GPU model and precision, and INT8 requires calibration data. The underlying mechanics map to the "inference engines" section of Inference: from forward pass to inference engine.

2. The Optimization Pipeline at a Glance ​

text
PyTorch model
   │  torch.onnx.export
   ▼
model.onnx
   │  trtexec --fp16 / --int8 (+ calibration)
   ▼
model.engine (bound to GPU model and TensorRT version)
   │  Triton backend / C++ Runtime / Python Runtime
   ▼
Runs on a Jetson or data-center GPU

3. Step 1: Export and Validate the ONNX ​

python
import torch
import onnx
from torchvision.models import resnet50, ResNet50_Weights

model = resnet50(weights=ResNet50_Weights.DEFAULT).eval()
dummy = torch.randn(1, 3, 224, 224)

torch.onnx.export(
    model, dummy, "resnet50.onnx",
    input_names=["input"], output_names=["output"],
    # Declare both dynamic batch and dynamic resolution explicitly (dynamic resolution is common on edge)
    dynamic_axes={"input": {0: "batch", 2: "height", 3: "width"},
                  "output": {0: "batch", 2: "height", 3: "width"}},
    opset_version=17,
)

# Validate the model structure with the onnx library
onnx.checker.check_model(onnx.load("resnet50.onnx"))
print("onnx ok, ops:", len(onnx.load("resnet50.onnx").graph.node))

Export essentials and compatibility checks are covered in model formats and conversion. When export fails or accuracy shifts, suspect custom operators and the opset first.

4. Step 2: Conversion with trtexec (FP16 / INT8 / Dynamic Shapes) ​

trtexec is TensorRT's built-in conversion and benchmarking tool (it ships with TensorRT):

bash
# FP16: the usual choice, zero extra cost, roughly 2x the speed
trtexec --onnx=resnet50.onnx --saveEngine=resnet50_fp16.engine \
        --fp16 \
        --minShapes=input:1x3x224x224 \
        --optShapes=input:1x3x224x224 \
        --maxShapes=input:16x3x224x224

# INT8: requires a calibration set (see Section 5), roughly 3-4x the speed
trtexec --onnx=resnet50.onnx --saveEngine=resnet50_int8.engine \
        --int8 \
        --calib=/data/calib/calib.txt \
        --calibCache=calib.cache \
        --minShapes=input:1x3x224x224 \
        --optShapes=input:1x3x224x224 \
        --maxShapes=input:16x3x224x224

The three dynamic-shape parameters, verdict first:

  • --minShapes / --optShapes / --maxShapes define the batch/resolution range: min sets the floor for memory preallocation, opt is the tuning target for kernel selection (the shape you'll hit most often), and max caps memory. The wider the range, the slower the build and the more memory reserved — keep it close to real usage;
  • --calibCache: caches calibration results to a file so the second build doesn't redo calibration (otherwise every --int8 build is slow);
  • The .engine artifact is bound to a specific GPU model and TensorRT version: an engine built on an A10 won't run on a Jetson; changing hardware means rebuilding.

5. INT8 Calibration and Accuracy Validation ​

INT8 quantization needs a representative calibration set so TensorRT can measure per-layer activation distributions and derive scale factors. Straight to the point:

  1. The calibration set must represent real production input — use 500–2000 real images sampled from production, not a random slice of the training set (its distribution may have drifted from production);
  2. Always validate accuracy after calibration: run a validation set and compare metrics against FP32/FP16 (Top-1 or your task metric). If INT8 drops more than your threshold (e.g. >0.5% for image classification), switch to per-layer INT8+FP16 mixing (--layer-precision) and only quantize the insensitive layers;
  3. Accuracy vs. speed table (ResNet50, Jetson Orin Nano):
PrecisionLatency (ms)Relative speedTop-1 delta
FP32~201.0×baseline
FP16~8~2.5×lossless
INT8~4.5~4.4×approx. -0.2% to -0.5%

For the general quantization trade-off (not just TensorRT), see quantization.

Per-Layer Precision Control: Plan B When INT8 Hurts Too Much ​

When whole-network INT8 degrades accuracy unacceptably, fall back to per-layer mixed precision: use --layer-precision to keep specific layers in FP16. Example command:

bash
trtexec --onnx=resnet50.onnx --saveEngine=resnet50_mixed.engine \
        --int8 --fp16 \
        --calib=/data/calib/calib.txt \
        # Keep sensitive layers (e.g. the final fully-connected and normalization layers) in FP16
        --layer-precision=fc:PFP16 \
        --layer-precision=conv_5: PFP16

Verdict first: try whole-network INT8 and measure the drop; if it's small, ship it; if it's large, tune per layer; if that still fails, consider QAT (quantization-aware training). Most CNNs lose <0.5% with whole-network INT8; Transformers and detection models concentrate sensitivity in specific layers, so per-layer control pays off visibly. You can dump the layer name list first with trtexec --dumpLayerInfo.

6. Runtime Inference: Loading the Engine with Python / C++ ​

python
# trt_runtime.py
import tensorrt as trt
import numpy as np
from cuda import cudart  # or pycuda

logger = trt.Logger(trt.Logger.WARNING)
runtime = trt.Runtime(logger)

# Deserialize the engine (binary; bound to the GPU it was built on)
with open("resnet50_fp16.engine", "rb") as f:
    engine = runtime.deserialize_cuda_engine(f.read())
context = engine.create_execution_context()

# Dynamic shapes: the input shape must be set explicitly before each inference
context.set_input_shape("input", (1, 3, 224, 224))

# Allocate input/output device buffers (flow only; real code must handle streams and synchronization)
input_buf, output_buf, stream = allocate_buffers(engine, context)
# Copy input in -> context.execute_async_v2(bindings, stream) -> copy output out

Key point: engines built with dynamic shapes require set_input_shape and buffer reallocation on every inference; in production you'll usually let Triton's TensorRT backend host the engine and skip hand-written memory management — it integrates seamlessly with Triton. The C++ path works the same way via InferRuntime; Jetson edge deployments commonly use C++, or Python + TensorRT from the Jetson-specific images.

Hosting the Engine with Triton: No Hand-Written Memory Management ​

When you run multiple engines in the data center or on a Jetson, Triton's tensorrt backend is the most trouble-free way to host them; just switch the backend to TensorRT in config.pbtxt:

protobuf
# models/resnet50/config.pbtxt
name: "resnet50"
backend: "tensorrt"
max_batch_size: 16
input  [ { name: "input",  data_type: TYPE_FP32, dims: [3, 224, 224] } ]
output [ { name: "output", data_type: TYPE_FP32, dims: [1000] } ]
instance_group [ { kind: KIND_GPU, count: 1 } ]

Benefits: dynamic batching, concurrent scheduling, and metrics output all come for free, and engine file management is folded into the model repository. It integrates seamlessly with NVIDIA Triton multi-model serving; Triton parameters such as --backend-config=tensorrt,coalesce-request-input=true let you tune batching further.

7. Edge Deployment: Jetson ​

Jetson is NVIDIA's edge computing platform. A comparison of common models (note the unified memory):

ModelComputeUnified memoryPowerTypical use
Jetson Orin Nano~40 TOPS (INT8)8/16GB7–25WReal-time detection at the camera
Jetson Orin NX~100 TOPS (INT8)16GB10–40WMulti-stream video at the edge
Jetson AGX Orin~275 TOPS (INT8)64GB15–60WEdge LLMs / heavy models

Three deployment essentials:

  1. Use Jetson-specific images: nvcr.io/nvidia/l4t-pytorch, l4t-tensorrt, or the JetPack SDK (which bundles a matching TensorRT version) — do not install data-center TensorRT on a Jetson; the architecture doesn't match;
  2. Cross-compilation vs. on-device builds: the engine must be rebuilt on the target Jetson (or built offline with trtexec on the board), and container images must be built for aarch64;
  3. Power and thermal constraints: Jetson has power modes (nvpmodel -m 0 for max performance / low-power modes); pick the right mode for heavy models and monitor temperatures — thermal throttling causes latency spikes. For hardware selection and power trade-offs, see GPU and hardware selection.

The Real Challenges of Edge Deployment

Edge is not the cloud: the network can drop, devices can reboot, and nobody is standing by to restart the process. Engine files must be persisted locally (avoid recompiling at every boot — a single build can take minutes); the service needs to auto-start on boot and be restarted after crashes; and model updates should arrive via OTA rather than on-site reflashing. These "operational costs" are easy to overlook, yet they dominate edge work.

8. Measured Performance Comparison ​

Same Jetson Orin Nano, ResNet50, batch size 1, dynamic input 224×224:

SetupLatency (ms)Throughput (FPS)Memory usage
PyTorch eager (FP32)~22~45~500MB
TensorRT FP16~8~125~300MB
TensorRT INT8~4.5~220~250MB

Conclusion: FP16 is the best-value starting point; move to INT8 once accuracy is acceptable. For mobile (no NVIDIA GPU), the alternative route is TFLite + quantization via TFLite/ML Kit; the strategy is the same (calibration + accuracy validation).

Common Pitfalls and Troubleshooting ​

PitfallSymptomFix
Engine fails on a different GPUengine file is generated on an incompatible deviceEngines are bound to GPU/version; rebuild on the target device
Dynamic shape not setInference throws invalid shapesCall set_input_shape before every inference
INT8 accuracy collapseTop-1 drops 5%+Calibration set not representative; switch to per-layer mixed precision
Conversion failstrtexec reports unsupported operatorsMove to a higher opset, replace custom operators, add --onnx-normalization
Device memory OOMDynamic shape range set too wideNarrow min/maxShapes; don't pad the max
Jetson can't keep upLatency far above expectationsCheck the nvpmodel power mode and thermal throttling
Version mismatchRuntime reports missing trt version APIUse the same TensorRT version for the engine and the runtime

Further Reading ​

References ​