Appearance
TensorRT and GPU Inference
1. Definition: Polishing a "General-Purpose Graph" into a "Blazing-Fast GPU Graph"
TensorRT is the GPU inference optimization engine (inference accelerator) NVIDIA has shipped since 2015. Its input is the "general-purpose computation graph" trained in PyTorch / TensorFlow / MXNet; its output is a Plan file bound to a specific GPU architecture and optimized layer by layer — think of it as a native executable compiled for this exact card, this batch size, and this precision.
PyTorch model (.pt)
│ tracing / scripting
▼
ONNX / TorchScript intermediate representation
│ TensorRT Builder (graph optimization + kernel auto-tuning + quantization)
▼
Plan file (.plan/.engine) ← bound to GPU architecture + precision + shape range
│ TensorRT Runtime
▼
Inference on the GPUThe key difference between TensorRT and PyTorch is "compilation vs. interpretation": PyTorch defaults to eager mode (executing line by line), while TensorRT is AOT (ahead-of-time) compilation — at build time it looks at the entire graph, fuses operators, picks the fastest kernels, and hard-codes the optimal execution plan. On classic models like ResNet-50 / BERT, it typically delivers a 2–5× latency reduction over PyTorch eager; on LLMs, combined with TensorRT-LLM (see TensorRT-LLM), it pushes even further.
To understand where TensorRT sits, look at it inside the inference stack (see Learning Paths: Three Routes and GPU Architecture and Optimization): training frameworks on top, the CUDA Driver below, and TensorRT in the middle acting as a "framework-agnostic GPU inference compiler."
2. Core Mechanism: Builder → Engine → Runtime
The TensorRT workflow has three stages. Understand these three, and you understand the entire mental model of TensorRT.
1. Builder: The Build-Time Optimization "Workshop"
The Builder takes a network definition and does four things:
| Technique | What it does | Typical gain |
|---|---|---|
| Graph optimization | Operator fusion (Conv+BN+ReLU, Linear+ReLU), constant folding, dead-code elimination | Fewer memory accesses, lower kernel launch overhead |
| Kernel auto-tuning | For each operator, picks several candidates from the kernel library and benchmarks them on the target GPU | The same operator picks different implementations for different batch/shape |
| Precision calibration | FP32 → FP16 / INT8 / FP8, with calibration or quantization-aware training | 2–4× throughput |
| Workspace planning | Allocates temporary GPU workspace for each operator | Trading space for speed |
For graph optimization details, see Computation Graph Optimization and Kernel Fusion and Custom Kernels. A typical fusion:
Original graph: MatMul → Bias → ReLU → MatMul → Bias → ReLU
Optimized: [FusedLinearRelu] → [FusedLinearRelu] (two kernel launches)2. Engine: The Build Artifact
The binary produced by the Builder is called the Engine (serialized into a .plan / .engine file). An Engine is bound to a GPU architecture — an engine built on an A100 cannot simply run on an H100 (unless you use --useUnifiedCapture or rebuild). This is the most common pitfall for beginners: when the CI build machine and the deployment machine have different GPU models, you must rebuild on the deployment machine, or unify the environment with containers.
Engines cannot migrate across architectures
Do not push a plan built on a local RTX 4090 straight to an A100 server. Common practice:
- Build on a card of the same model as production and ship the plan file with the image;
- Build on the deployment machine with
trtexec --saveEngine; - Build inside a container and distribute it with the image (recommended, but the driver/CUDA versions inside the image must align).
3. Runtime: Execution Time
The Runtime loads the Engine, creates an ExecutionContext, binds input/output buffers, and runs inference via enqueueV2 / execute_v3. A Context is a stateful object — the same engine can create multiple contexts running in parallel (subject to GPU SM resource constraints), which underpins the "concurrent model execution" in Triton (see Triton Inference Server).
python
import tensorrt as trt
import pycuda.driver as cuda # or cuda-python
runtime = trt.Runtime(trt.Logger())
with open("resnet50.engine", "rb") as f:
engine = runtime.deserialize_cuda_engine(f.read())
context = engine.create_execution_context()
# Allocate input/output device buffers; binding maps name → buffer address3. INT8 Quantization and Calibration
INT8 is one of TensorRT's signature capabilities. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision; here we focus on TensorRT's implementation path.
PTQ Path: Entropy Calibration
TensorRT performs PTQ (post-training quantization) with the IInt8EntropyCalibrator2:
- Prepare a few hundred representative inputs (calibration dataset);
- The Builder runs these inputs in FP32 and collects the activation distribution of every layer;
- For each layer, it picks a
(scale, zero_point)pair that minimizes the KL divergence before and after quantization (hence "entropy"); - The scale is written into the engine, and inference runs in INT8.
python
import tensorrt as trt
# Custom calibrator
class MyCalibrator(trt.IInt8EntropyCalibrator2):
def __init__(self, data_loader, cache_file="int8.cache"):
super().__init__()
self.data_loader = data_loader
self.cache_file = cache_file
self.batch_iter = iter(data_loader)
self.device_input = cuda.mem_alloc(...) # allocate per batch size
def get_batch(self, names):
try:
batch = next(self.batch_iter)
except StopIteration:
return None
cuda.memcpy_htod(self.device_input, batch.numpy())
return [int(self.device_input)]
def read_calibration_cache(self):
if os.path.exists(self.cache_file):
with open(self.cache_file, "rb") as f:
return f.read()
def write_calibration_cache(self, cache):
with open(self.cache_file, "wb") as f:
f.write(cache)
builder.int8_calibrator = MyCalibrator(...)
builder.config.set_flag(trt.BuilderFlag.INT8)QAT Path: Importing from PyTorch Quantization-Aware Training
If PTQ degrades accuracy badly (even <1% counts as serious), go QAT: train a model with fake-quantization operators in PyTorch / the pytorch-quantization toolkit, export it to ONNX with QuantizeLinear/DequantizeLinear nodes, and TensorRT 8.x+ will recognize them and generate an INT8 engine directly. QAT is typically 0.3–1 points more accurate than PTQ, at a higher training cost.
The Precision Ladder
| Precision | GPU memory | Throughput | Accuracy loss | When to use |
|---|---|---|---|---|
| FP32 | 1× | 1× | 0 | Validation baseline |
| FP16 | 0.5× | ~2× | <0.1% | Almost always on |
| INT8 (PTQ) | 0.25× | 3–4× | 0.3–1% | Mainstream for detection/classification |
| INT8 (QAT) | 0.25× | 3–4× | <0.3% | Accuracy-sensitive scenarios |
| FP8 (H100+) | 0.25× | 4–6× | <0.5% | The new favorite for large models |
4. The Plugin Interface: The Escape Hatch for Custom Operators
What if the model contains operators TensorRT doesn't know? Two paths:
- Rewrite the model: replace them with an equivalent combination of TensorRT-supported operators (recommended, better maintainability).
- Write a Plugin: implement the
IPluginV2/IPluginV2DynamicExt/IPluginV3interfaces, register them in the registry, and TensorRT resolves them by name at build time.
cpp
// Simplified Plugin interface
class CustomLayerPlugin : public nvinfer1::IPluginV2DynamicExt {
const char* getPluginType() const noexcept override { return "CustomLayer"; }
int enqueue(...) noexcept override {
// call a custom CUDA kernel
custom_layer_cuda_forward(...);
return 0;
}
// ... clone / serialize / getOutputDimensions etc.
};
REGISTER_TENSORRT_PLUGIN(CustomLayerPlugin);Plugins let TensorRT both "optimize standard operators" and "escape to do non-standard things" — the last line of defense for compatibility. The cost: plugin operators do not participate in graph optimization (they cannot fuse with neighboring operators) and usually run slower than native ones.
The Plugin trade-off
If you can express it with standard operators, don't write a Plugin. Maintenance is costly, the build pipeline gets complicated, and cross-version APIs are unstable (V2 → V2DynamicExt → V3 went through several breaking changes). But some custom attention / projection layers genuinely require a Plugin — mastering it remains an engineering must.
5. Relationship with PyTorch: torch_tensorrt
Going through ONNX → TensorRT directly is tedious: export ONNX, check opset support, tune plugins. torch_tensorrt is NVIDIA's "direct-connect" solution — it takes a torch.nn.Module or torch.export graph directly, compiles it to TensorRT internally with Dynamo/TorchScript, and falls back to PyTorch eager for uncovered operators (hybrid graph).
python
import torch_tensorrt
model = model.eval().cuda()
inputs = [torch.randn(1, 3, 224, 224).cuda()]
trt_model = torch_tensorrt.compile(
model,
ir="torch_compile", # new pipeline, based on Dynamo
inputs=inputs,
enabled_precisions={torch.float, torch.half}, # enable FP16
workspace_size=1 << 30,
min_block_size=5, # minimum subgraph size; smaller ones fall back to eager
)
# afterwards trt_model is called just like the original model
out = trt_model(inputs[0])Since torch_tensorrt 1.x, ir="torch_compile" has been the recommended path, aligned with the PyTorch 2.x torch.compile ecosystem — the de facto standard for PyTorch inference acceleration after 2024. See Computation Graph Optimization.
6. Performance Data: Baseline Reference
Here is a set of baseline numbers for "order-of-magnitude" reference only — benchmark with your own model + card (methodology in Inference Benchmarking in Practice):
| Model | Hardware | Precision | Latency (batch=1) | vs. PyTorch eager |
|---|---|---|---|---|
| ResNet-50 | A100 80G | FP16 | ~0.4 ms | 3–4× |
| BERT-base | A100 80G | FP16 | ~1.0 ms | 3× |
| ResNet-50 | T4 | INT8 | ~1.0 ms | 5–6× |
| YOLOv8-m | A100 | FP16 | ~2.5 ms | 3× |
| ViT-B/16 | A100 | FP16 | ~2.0 ms | 3–4× |
Pitfalls of baseline numbers
Any "TensorRT is X times faster than PyTorch" number is affected by the following:
- Whether the PyTorch baseline enables AMP /
torch.compile: withtorch.compile, the gap usually shrinks to 1.3–1.8×; - Batch size: the gap narrows at large batch (the bottleneck shifts to memory bandwidth);
- Dynamic shapes: dynamic shapes compress TensorRT's optimization space;
- Warmup: the first inference includes CUDA context initialization and must be excluded.
7. Limitations and Boundaries
TensorRT is not a silver bullet. Its "boundaries" come from design trade-offs:
- GPU architecture binding: Engines cannot migrate across architectures, so CI/CD pipelines must build per target card; FPGA, AMD GPU, and Apple Silicon are entirely unsupported — for those scenarios, consider ONNX Runtime or OpenVINO.
- Dynamic shapes are painful: supported, but you must pre-declare an
OptProfile(min/opt/max); going out of range either errors out or performance collapses. For LLMs, where seq_len is highly dynamic, TensorRT-LLM solves this with "Paged KV Cache + In-Flight Batching" (see TensorRT-LLM). - Incomplete operator coverage: models with new operators or non-standard structures (e.g., certain attention variants) easily hit "unsupported op" — either write a Plugin or change the model.
- Long build times: the Builder's kernel auto-tuning is expensive; building a large model takes anywhere from hours to tens of hours. CI must cache plans.
- Closed source: open-source alternatives (TVM, MLC, TensorRT's open-source
ONNX Parser) still lag in maturity. - Steep learning curve: the API broke multiple times across 7.x/8.x/8.6/10.x, and the docs-to-version mapping is complex.
8. Relationship with Similar Solutions
| Solution | Positioning | Relationship |
|---|---|---|
| ONNX Runtime | Cross-platform general-purpose engine (see ONNX Runtime) | ORT has a TensorRT Execution Provider — ORT is the upper layer, TensorRT is one of its backends |
| OpenVINO | Intel CPU/iGPU optimization (see OpenVINO) | The same "graph optimization + quantization" playbook, but targeting different hardware |
| torch.compile / inductor | PyTorch-native AOT compilation | Two paths — "native vs. flagship"; torch_tensorrt bridges them |
| TensorRT-LLM | LLM-specific extension | Builds PagedAttention, In-Flight Batching, continuous batching, and other LLM optimizations on top of TensorRT (see TensorRT-LLM) |
| vLLM / SGLang | LLM inference engines (see vLLM) | Complementary rather than competing — TensorRT-LLM is NVIDIA's counterpart to vLLM |
9. Where to Go Next
- Concept pages: Computation Graph Optimization, Kernel Fusion and Custom Kernels, Model Quantization Fundamentals, Weight-Only Quantization and Mixed Precision, GPU Architecture and Optimization, Latency, Throughput, and Concurrency
- Case-study pages: ONNX Runtime, OpenVINO, TensorRT-LLM, Triton Inference Server, Distributed Inference (TP/PP)
- Practice pages: Inference Engine Comparison, Tuning and Performance Optimization, Inference Benchmarking in Practice, Common Pitfalls and Anti-Patterns
- Papers and hardware: Classic Papers in Depth, Hardware Primer
References
- NVIDIA. TensorRT Developer Guide — official documentation (align with your version)
- NVIDIA. TensorRT Sample Plugins — open-source repo (includes parser and plugin examples)
- NVIDIA. torch_tensorrt — official PyTorch integration
- NVIDIA. trtexec — command-line build/bench tool
- Migacz. 8-bit Inference with TensorRT (GTC 2017) — classic lecture on the INT8 calibrator
- NVIDIA. FP8 Quantization with TensorRT (2023) — FP8 practice on H100