Skip to content

TensorRT-LLM

At a glance NVIDIA's LLM-specific extension on top of TensorRT, integrating In-Flight Batching, Paged KV Cache, FlashAttention, FP8, and more — the flagship stack for LLM inference on H100. This article dissects the build flow, tuning switches, the comparison with vLLM, and Triton deployment.

TensorRT-LLM ​

1. Definition: NVIDIA's "Official Flagship" for LLM Inference ​

TensorRT-LLM is the LLM inference optimization library NVIDIA released in 2023, built on top of TensorRT and deeply optimized for large language models (decoder-only Transformers). Its positioning can be summed up in one sentence: the stack NVIDIA's own engineers squeezed to the limit, on NVIDIA's own GPUs, for NVIDIA's own LLM inference.

To understand its place, look at the full NVIDIA inference stack:

                ┌─────────────────────────────────────┐
                │   Application / serving (OpenAI     │
                │   protocol, APIs)                   │
                ├─────────────────────────────────────┤
                │  Triton LLM Backend (serving layer) │  ← see Triton Inference Server
                ├─────────────────────────────────────┤
                │  TensorRT-LLM (LLM optimization)    │  ← the topic of this page
                │   ├─ In-Flight Batching             │
                │   ├─ Paged KV Cache                 │
                │   ├─ FlashAttention / FP8           │
                │   └─ SmoothQuant / GPTQ / AWQ       │
                ├─────────────────────────────────────┤
                │  TensorRT (graph / kernel layer)    │  ← see the TensorRT case study
                ├─────────────────────────────────────┤
                │  CUDA Driver / cuBLAS / cuDNN       │
                └─────────────────────────────────────┘

TensorRT-LLM and vLLM are direct competitors — vLLM is the flagship of the open-source community, TRT-LLM the flagship of NVIDIA. On H100, TRT-LLM is usually slightly faster (1.1–1.3×), but vLLM wins on ease of use, community ecosystem, and out-of-the-box HuggingFace model support.

2. Core Features ​

1. In-Flight Batching (Same as vLLM's Continuous Batching) ​

NVIDIA renamed continuous batching to In-Flight Batching (also called dynamic batching / iteration-level batching) — the same idea as vLLM:

  • Mixed prefill/decode scheduling: every step can admit new requests and retire finished ones.
  • Chunked prefill: long prompts are split into chunks that run together with decode, so a long prompt cannot starve short requests.
  • Priority scheduling: request priorities can be set so high-priority requests cut in line.

See Batching and Request Scheduling.

2. Paged KV Cache ​

TRT-LLM also has a paged KV cache (same idea as vLLM's PagedAttention, but an independent implementation). The tensor memory pool on each H100 is managed in blocks: requests allocate on demand, and physical memory stays compact.

The story behind the name

PagedAttention is the name from the vLLM paper, but the idea of paged KV caches is not original to vLLM — the OS-virtual-memory analogy already appeared in 2022 papers. TRT-LLM and vLLM "rediscovered the same principle in parallel"; the implementation details differ, and TRT-LLM's paged KV cache is more tightly integrated with FlashAttention / FP8.

3. FlashAttention Integration ​

TRT-LLM enables FlashAttention 2/3 by default (FA3 on Hopper), reducing the IO complexity of attention from O(n²) to O(n²/M) (where M is the SRAM size). For FlashAttention details, see Classic Papers in Depth and GPU Architecture and Optimization.

Traditional attention:  S = Q·K^T  (n×n matrix written to HBM)  →  P = softmax(S)  →  O = P·V
                        HBM reads/writes: O(n²)

FlashAttention: split Q/K/V into tiles, complete attention inside SRAM, no intermediate matrix written
                        HBM reads/writes: O(n²/M)

4. FP8: The H100 Killer Feature ​

The H100 Hopper architecture natively supports FP8 (E4M3 / E5M2), with twice the theoretical throughput of BF16. TRT-LLM supports:

  • FP8 weights + FP8 activations (W8A8): extreme throughput
  • FP8 weights + BF16 activations (W8A16): accuracy first
  • FP8 KV cache: halves KV cache memory — a must for long-context scenarios
bash
# Quantize Llama-3 to FP8 with the NeMo framework
python examples/llama/quantize.py \
    --model_dir /models/Meta-Llama-3-8B \
    --dtype float16 \
    --qformat fp8 \
    --output_dir /models/llama3-8b-fp8

5. Quantization Support Matrix ​

MethodBit widthTRT-LLM supportNotes
SmoothQuantINT8✅Activations smoothed, W8A8, near-lossless
GPTQINT4✅Paired with Marlin kernels
AWQINT4✅Same as vLLM
FP8FP8✅Recommended on H100
INT4 AWQ + FP8 KVINT4 + FP8✅First choice for long context
INT8 KV cacheINT8✅Long context on non-H100 cards

See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

3. Build Flow: Model → Engine ​

TRT-LLM splits the build into two steps: first quantize / convert, then build the engine.

bash
# Step 1: convert + quantize (NeMo or TRT-LLM's native tools)
python tools/llm/llama/convert.py \
    --model_dir /models/Meta-Llama-3-8B \
    --output_dir /models/llama3-8b-trt-checkpoint \
    --dtype float16 \
    --use_smooth_quant        # or --quantize_fp8 / --use_gptq / --use_awq

# Step 2: build the engine
trtllm-build \
    --checkpoint_dir /models/llama3-8b-trt-checkpoint \
    --output_dir /models/llama3-8b-engine \
    --gemm_plugin float16 \
    --gpt_attention_plugin float16 \
    --context_fmha enable \
    --paged_kv_cache enable \
    --remove_input_padding enable \
    --use_paged_context_fmha enable \
    --max_batch_size 32 \
    --max_input_len 2048 \
    --max_output_len 512 \
    --max_num_tokens 8192 \
    --builder_optimization_level 3

Key Build Switches ​

SwitchPurposeRecommendation
--gemm_pluginUse TRT-LLM's own GEMM (replacing TRT's default)float16 / bfloat16 / fp8
--gpt_attention_pluginUse the FlashAttention variantfloat16 / fp8
--context_fmhaFlashAttention in the prefill phaseenable (always on)
--use_paged_context_fmhaPaged KV in prefill as wellRecommended on H100
--paged_kv_cachePaged KV cacheenable (always on)
--remove_input_paddingRemove padding, packed sequencesenable (always on)
--max_num_tokensMax tokens per stepDetermines the batch × seq ceiling
--builder_optimization_levelOptimization depth (0–5)3–5, trading build time for performance

Builds take a long time

--builder_optimization_level 5 + a large model → a single build takes hours to tens of hours. CI must cache engines (named by model hash) and never rebuild every time. Same as TensorRT.

4. Performance Tuning: Measured Paths ​

Here is a set of baselines on H100 80GB (see Inference Benchmarking in Practice):

ModelPrecisionConcurrencyThroughput (tokens/s)First-token latency
Llama-3-8BBF1632~550028 ms
Llama-3-8BFP8 (W8A8)32~1000028 ms
Llama-3-8BINT4 AWQ + FP8 KV32~1200030 ms
Llama-3-70BBF16 (2×H100 TP)32~170075 ms
Llama-3-70BFP8 (1×H100)32~190070 ms
Mixtral-8x22BFP8 (4×H100 TP)32~2000120 ms

Against vLLM 0.6.x (same hardware):

Llama-3-8B FP8:        TRT-LLM  ~10000 tokens/s    vs  vLLM  ~8500 tokens/s   (1.18×)
Llama-3-70B FP8 TP=2:  TRT-LLM  ~1900  tokens/s    vs  vLLM  ~1700  tokens/s   (1.12×)

TRT-LLM is slightly faster on H100; on A100 / L40 and similar cards the gap narrows or even reverses. Don't assume the outcome — benchmark first.

5. Tuning in Practice: Common Switches ​

1. Long-Context Optimization ​

bash
# Llama-3.1-8B-Instruct with 128K context
trtllm-build \
    ... \
    --max_input_len 127000 \
    --max_output_len 1000 \
    --max_num_tokens 131072 \
    --context_fmha enable \
    --use_paged_context_fmha enable \
    --paged_kv_cache enable \
    --kv_cache_quant_type INT8    # or FP8, halving the KV cache

The long-context bottleneck is KV cache memory. The KV cache of Llama-3-8B at 128K context is ~40 GB, or ~20 GB after FP8 quantization — fitting on a single card.

2. Speculative Decoding ​

TRT-LLM 0.10+ supports speculative decoding (see Speculative Decoding and Medusa/EAGLE):

bash
# EAGLE speculative decoding
trtllm-build \
    ... \
    --speculative_decoding_mode eagle \
    --eagle_model_dir /models/llama3-8b-eagle

EAGLE-3 achieves 2–4× on SGLang/vLLM; TRT-LLM is similar.

3. Triton Deployment ​

The official deployment form of TRT-LLM is the Triton LLM Backend (see Triton Inference Server):

models/
└── llama3-8b/
    ├── 1/
    │   └── (path to the TRT-LLM engine)
    └── config.pbtxt

config.pbtxt:

backend: "tensorrtllm"
max_batch_size: 32

input [
  { name: "input_ids",     data_type: TYPE_INT32,  dims: [-1] },
  { name: "input_lengths", data_type: TYPE_INT32,  dims: [1] },
  { name: "request_output_len", data_type: TYPE_INT32, dims: [-1] },
  { name: "stop_words_list",    data_type: TYPE_INT32, dims: [-1] }
]
output [
  { name: "output_ids",        data_type: TYPE_INT32, dims: [-1, -1] },
  { name: "output_lengths",    data_type: TYPE_INT32, dims: [-1] },
  { name: "cum_log_probs",     data_type: TYPE_FP32,  dims: [-1] }
]

parameters: {
  key: "decoder_model_path"
  value: { string_value: "/models/llama3-8b/engine" }
}
parameters: {
  key: "max_beam_width"
  value: { string_value: "1" }
}
parameters: {
  key: "TRTLLM_Batch_scheduler_policy"
  value: { string_value: "max_utilization" }
}

6. Comparison with vLLM ​

DimensionTensorRT-LLMvLLM
H100 performance1.0–1.3×, slightly aheadThe baseline
A100 performanceComparableComparable
Ease of useHeavy build flow, steep learning curveOne line to start a server
Model supportMainstream models covered; new models wait for NVIDIA adaptationHuggingFace models out of the box
QuantizationFP8 / SmoothQuant / GPTQ / AWQAWQ / GPTQ / FP8 / bitsandbytes
CommunityNVIDIA-led, enterprise-heavyUC Berkeley-led, community-driven
Structured outputoutlines / xgrammar integrationThe SGLang family is stronger
MultimodalSupported but relatively behindDecent since 0.5+
Cross-hardwareNVIDIA onlyNVIDIA only (same)
Triton integrationNative LLM BackendVia the vLLM Triton backend

Decision guide

  • H100 cluster, chasing the extreme: try TRT-LLM first and keep it if benchmarks favor it; if the difference is small, use vLLM (ease-of-use advantage).
  • A100 / L40 / T4: vLLM is usually the less troublesome choice.
  • Quick PoCs, small teams: vLLM.
  • NVIDIA partners / enterprise production: the full TRT-LLM + Triton stack, backed by NVIDIA commercial support.
  • Structured output / agents: consider SGLang instead of TRT-LLM.

7. Limitations and Boundaries ​

  1. Steep learning curve: dozens of build switches and documentation scattered across GitHub issues; beginners give up easily.
  2. Long build times: kernel auto-tuning is extremely slow; CI must cache engines.
  3. Lagging support for new models: new architectures (Mamba, Hyena, new MoE variants) land 1–2 months behind vLLM.
  4. Bound to NVIDIA: unavailable for cross-hardware scenarios — go with ONNX Runtime or llama.cpp.
  5. Complex Triton configuration: a miswritten config.pbtxt is an easy trap; you need to be familiar with the Triton LLM backend protocol.
  6. Closed-source components: the core plugins are closed, making debugging hard.

8. Where to Go Next ​

References ​