Skip to content

TensorRT and GPU Inference

At a glance NVIDIA's own GPU inference optimization engine, which turns the "general-purpose compute" of training frameworks into "extreme throughput on the GPU" through graph fusion, kernel auto-tuning, and INT8/FP8 quantization. This article dissects the three-stage Builder/Engine/Runtime pipeline, INT8 calibration, the Plugin interface, and torch_tensorrt integration.

TensorRT and GPU Inference ​

1. Definition: Polishing a "General-Purpose Graph" into a "Blazing-Fast GPU Graph" ​

TensorRT is the GPU inference optimization engine (inference accelerator) NVIDIA has shipped since 2015. Its input is the "general-purpose computation graph" trained in PyTorch / TensorFlow / MXNet; its output is a Plan file bound to a specific GPU architecture and optimized layer by layer — think of it as a native executable compiled for this exact card, this batch size, and this precision.

PyTorch model (.pt)
      │  tracing / scripting
      ▼
ONNX / TorchScript intermediate representation
      │  TensorRT Builder (graph optimization + kernel auto-tuning + quantization)
      ▼
Plan file (.plan/.engine)  ← bound to GPU architecture + precision + shape range
      │  TensorRT Runtime
      ▼
Inference on the GPU

The key difference between TensorRT and PyTorch is "compilation vs. interpretation": PyTorch defaults to eager mode (executing line by line), while TensorRT is AOT (ahead-of-time) compilation — at build time it looks at the entire graph, fuses operators, picks the fastest kernels, and hard-codes the optimal execution plan. On classic models like ResNet-50 / BERT, it typically delivers a 2–5× latency reduction over PyTorch eager; on LLMs, combined with TensorRT-LLM (see TensorRT-LLM), it pushes even further.

To understand where TensorRT sits, look at it inside the inference stack (see Learning Paths: Three Routes and GPU Architecture and Optimization): training frameworks on top, the CUDA Driver below, and TensorRT in the middle acting as a "framework-agnostic GPU inference compiler."

2. Core Mechanism: Builder → Engine → Runtime ​

The TensorRT workflow has three stages. Understand these three, and you understand the entire mental model of TensorRT.

1. Builder: The Build-Time Optimization "Workshop" ​

The Builder takes a network definition and does four things:

TechniqueWhat it doesTypical gain
Graph optimizationOperator fusion (Conv+BN+ReLU, Linear+ReLU), constant folding, dead-code eliminationFewer memory accesses, lower kernel launch overhead
Kernel auto-tuningFor each operator, picks several candidates from the kernel library and benchmarks them on the target GPUThe same operator picks different implementations for different batch/shape
Precision calibrationFP32 → FP16 / INT8 / FP8, with calibration or quantization-aware training2–4× throughput
Workspace planningAllocates temporary GPU workspace for each operatorTrading space for speed

For graph optimization details, see Computation Graph Optimization and Kernel Fusion and Custom Kernels. A typical fusion:

Original graph:   MatMul → Bias → ReLU → MatMul → Bias → ReLU
Optimized:        [FusedLinearRelu] → [FusedLinearRelu]   (two kernel launches)

2. Engine: The Build Artifact ​

The binary produced by the Builder is called the Engine (serialized into a .plan / .engine file). An Engine is bound to a GPU architecture — an engine built on an A100 cannot simply run on an H100 (unless you use --useUnifiedCapture or rebuild). This is the most common pitfall for beginners: when the CI build machine and the deployment machine have different GPU models, you must rebuild on the deployment machine, or unify the environment with containers.

Engines cannot migrate across architectures

Do not push a plan built on a local RTX 4090 straight to an A100 server. Common practice:

  1. Build on a card of the same model as production and ship the plan file with the image;
  2. Build on the deployment machine with trtexec --saveEngine;
  3. Build inside a container and distribute it with the image (recommended, but the driver/CUDA versions inside the image must align).

3. Runtime: Execution Time ​

The Runtime loads the Engine, creates an ExecutionContext, binds input/output buffers, and runs inference via enqueueV2 / execute_v3. A Context is a stateful object — the same engine can create multiple contexts running in parallel (subject to GPU SM resource constraints), which underpins the "concurrent model execution" in Triton (see Triton Inference Server).

python
import tensorrt as trt
import pycuda.driver as cuda  # or cuda-python

runtime = trt.Runtime(trt.Logger())
with open("resnet50.engine", "rb") as f:
    engine = runtime.deserialize_cuda_engine(f.read())

context = engine.create_execution_context()
# Allocate input/output device buffers; binding maps name → buffer address

3. INT8 Quantization and Calibration ​

INT8 is one of TensorRT's signature capabilities. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision; here we focus on TensorRT's implementation path.

PTQ Path: Entropy Calibration ​

TensorRT performs PTQ (post-training quantization) with the IInt8EntropyCalibrator2:

  1. Prepare a few hundred representative inputs (calibration dataset);
  2. The Builder runs these inputs in FP32 and collects the activation distribution of every layer;
  3. For each layer, it picks a (scale, zero_point) pair that minimizes the KL divergence before and after quantization (hence "entropy");
  4. The scale is written into the engine, and inference runs in INT8.
python
import tensorrt as trt

# Custom calibrator
class MyCalibrator(trt.IInt8EntropyCalibrator2):
    def __init__(self, data_loader, cache_file="int8.cache"):
        super().__init__()
        self.data_loader = data_loader
        self.cache_file = cache_file
        self.batch_iter = iter(data_loader)
        self.device_input = cuda.mem_alloc(...)  # allocate per batch size

    def get_batch(self, names):
        try:
            batch = next(self.batch_iter)
        except StopIteration:
            return None
        cuda.memcpy_htod(self.device_input, batch.numpy())
        return [int(self.device_input)]

    def read_calibration_cache(self):
        if os.path.exists(self.cache_file):
            with open(self.cache_file, "rb") as f:
                return f.read()

    def write_calibration_cache(self, cache):
        with open(self.cache_file, "wb") as f:
            f.write(cache)

builder.int8_calibrator = MyCalibrator(...)
builder.config.set_flag(trt.BuilderFlag.INT8)

QAT Path: Importing from PyTorch Quantization-Aware Training ​

If PTQ degrades accuracy badly (even <1% counts as serious), go QAT: train a model with fake-quantization operators in PyTorch / the pytorch-quantization toolkit, export it to ONNX with QuantizeLinear/DequantizeLinear nodes, and TensorRT 8.x+ will recognize them and generate an INT8 engine directly. QAT is typically 0.3–1 points more accurate than PTQ, at a higher training cost.

The Precision Ladder ​

PrecisionGPU memoryThroughputAccuracy lossWhen to use
FP321×1×0Validation baseline
FP160.5×~2×<0.1%Almost always on
INT8 (PTQ)0.25×3–4×0.3–1%Mainstream for detection/classification
INT8 (QAT)0.25×3–4×<0.3%Accuracy-sensitive scenarios
FP8 (H100+)0.25×4–6×<0.5%The new favorite for large models

4. The Plugin Interface: The Escape Hatch for Custom Operators ​

What if the model contains operators TensorRT doesn't know? Two paths:

  1. Rewrite the model: replace them with an equivalent combination of TensorRT-supported operators (recommended, better maintainability).
  2. Write a Plugin: implement the IPluginV2 / IPluginV2DynamicExt / IPluginV3 interfaces, register them in the registry, and TensorRT resolves them by name at build time.
cpp
// Simplified Plugin interface
class CustomLayerPlugin : public nvinfer1::IPluginV2DynamicExt {
    const char* getPluginType() const noexcept override { return "CustomLayer"; }
    int enqueue(...) noexcept override {
        // call a custom CUDA kernel
        custom_layer_cuda_forward(...);
        return 0;
    }
    // ... clone / serialize / getOutputDimensions etc.
};

REGISTER_TENSORRT_PLUGIN(CustomLayerPlugin);

Plugins let TensorRT both "optimize standard operators" and "escape to do non-standard things" — the last line of defense for compatibility. The cost: plugin operators do not participate in graph optimization (they cannot fuse with neighboring operators) and usually run slower than native ones.

The Plugin trade-off

If you can express it with standard operators, don't write a Plugin. Maintenance is costly, the build pipeline gets complicated, and cross-version APIs are unstable (V2 → V2DynamicExt → V3 went through several breaking changes). But some custom attention / projection layers genuinely require a Plugin — mastering it remains an engineering must.

5. Relationship with PyTorch: torch_tensorrt ​

Going through ONNX → TensorRT directly is tedious: export ONNX, check opset support, tune plugins. torch_tensorrt is NVIDIA's "direct-connect" solution — it takes a torch.nn.Module or torch.export graph directly, compiles it to TensorRT internally with Dynamo/TorchScript, and falls back to PyTorch eager for uncovered operators (hybrid graph).

python
import torch_tensorrt

model = model.eval().cuda()
inputs = [torch.randn(1, 3, 224, 224).cuda()]

trt_model = torch_tensorrt.compile(
    model,
    ir="torch_compile",            # new pipeline, based on Dynamo
    inputs=inputs,
    enabled_precisions={torch.float, torch.half},  # enable FP16
    workspace_size=1 << 30,
    min_block_size=5,              # minimum subgraph size; smaller ones fall back to eager
)

# afterwards trt_model is called just like the original model
out = trt_model(inputs[0])

Since torch_tensorrt 1.x, ir="torch_compile" has been the recommended path, aligned with the PyTorch 2.x torch.compile ecosystem — the de facto standard for PyTorch inference acceleration after 2024. See Computation Graph Optimization.

6. Performance Data: Baseline Reference ​

Here is a set of baseline numbers for "order-of-magnitude" reference only — benchmark with your own model + card (methodology in Inference Benchmarking in Practice):

ModelHardwarePrecisionLatency (batch=1)vs. PyTorch eager
ResNet-50A100 80GFP16~0.4 ms3–4×
BERT-baseA100 80GFP16~1.0 ms3×
ResNet-50T4INT8~1.0 ms5–6×
YOLOv8-mA100FP16~2.5 ms3×
ViT-B/16A100FP16~2.0 ms3–4×

Pitfalls of baseline numbers

Any "TensorRT is X times faster than PyTorch" number is affected by the following:

  1. Whether the PyTorch baseline enables AMP / torch.compile: with torch.compile, the gap usually shrinks to 1.3–1.8×;
  2. Batch size: the gap narrows at large batch (the bottleneck shifts to memory bandwidth);
  3. Dynamic shapes: dynamic shapes compress TensorRT's optimization space;
  4. Warmup: the first inference includes CUDA context initialization and must be excluded.

7. Limitations and Boundaries ​

TensorRT is not a silver bullet. Its "boundaries" come from design trade-offs:

  1. GPU architecture binding: Engines cannot migrate across architectures, so CI/CD pipelines must build per target card; FPGA, AMD GPU, and Apple Silicon are entirely unsupported — for those scenarios, consider ONNX Runtime or OpenVINO.
  2. Dynamic shapes are painful: supported, but you must pre-declare an OptProfile (min/opt/max); going out of range either errors out or performance collapses. For LLMs, where seq_len is highly dynamic, TensorRT-LLM solves this with "Paged KV Cache + In-Flight Batching" (see TensorRT-LLM).
  3. Incomplete operator coverage: models with new operators or non-standard structures (e.g., certain attention variants) easily hit "unsupported op" — either write a Plugin or change the model.
  4. Long build times: the Builder's kernel auto-tuning is expensive; building a large model takes anywhere from hours to tens of hours. CI must cache plans.
  5. Closed source: open-source alternatives (TVM, MLC, TensorRT's open-source ONNX Parser) still lag in maturity.
  6. Steep learning curve: the API broke multiple times across 7.x/8.x/8.6/10.x, and the docs-to-version mapping is complex.

8. Relationship with Similar Solutions ​

SolutionPositioningRelationship
ONNX RuntimeCross-platform general-purpose engine (see ONNX Runtime)ORT has a TensorRT Execution Provider — ORT is the upper layer, TensorRT is one of its backends
OpenVINOIntel CPU/iGPU optimization (see OpenVINO)The same "graph optimization + quantization" playbook, but targeting different hardware
torch.compile / inductorPyTorch-native AOT compilationTwo paths — "native vs. flagship"; torch_tensorrt bridges them
TensorRT-LLMLLM-specific extensionBuilds PagedAttention, In-Flight Batching, continuous batching, and other LLM optimizations on top of TensorRT (see TensorRT-LLM)
vLLM / SGLangLLM inference engines (see vLLM)Complementary rather than competing — TensorRT-LLM is NVIDIA's counterpart to vLLM

9. Where to Go Next ​

References ​