Appearance
TensorRT-LLM
1. Definition: NVIDIA's "Official Flagship" for LLM Inference
TensorRT-LLM is the LLM inference optimization library NVIDIA released in 2023, built on top of TensorRT and deeply optimized for large language models (decoder-only Transformers). Its positioning can be summed up in one sentence: the stack NVIDIA's own engineers squeezed to the limit, on NVIDIA's own GPUs, for NVIDIA's own LLM inference.
To understand its place, look at the full NVIDIA inference stack:
┌─────────────────────────────────────┐
│ Application / serving (OpenAI │
│ protocol, APIs) │
├─────────────────────────────────────┤
│ Triton LLM Backend (serving layer) │ ← see Triton Inference Server
├─────────────────────────────────────┤
│ TensorRT-LLM (LLM optimization) │ ← the topic of this page
│ ├─ In-Flight Batching │
│ ├─ Paged KV Cache │
│ ├─ FlashAttention / FP8 │
│ └─ SmoothQuant / GPTQ / AWQ │
├─────────────────────────────────────┤
│ TensorRT (graph / kernel layer) │ ← see the TensorRT case study
├─────────────────────────────────────┤
│ CUDA Driver / cuBLAS / cuDNN │
└─────────────────────────────────────┘TensorRT-LLM and vLLM are direct competitors — vLLM is the flagship of the open-source community, TRT-LLM the flagship of NVIDIA. On H100, TRT-LLM is usually slightly faster (1.1–1.3×), but vLLM wins on ease of use, community ecosystem, and out-of-the-box HuggingFace model support.
2. Core Features
1. In-Flight Batching (Same as vLLM's Continuous Batching)
NVIDIA renamed continuous batching to In-Flight Batching (also called dynamic batching / iteration-level batching) — the same idea as vLLM:
- Mixed prefill/decode scheduling: every step can admit new requests and retire finished ones.
- Chunked prefill: long prompts are split into chunks that run together with decode, so a long prompt cannot starve short requests.
- Priority scheduling: request priorities can be set so high-priority requests cut in line.
See Batching and Request Scheduling.
2. Paged KV Cache
TRT-LLM also has a paged KV cache (same idea as vLLM's PagedAttention, but an independent implementation). The tensor memory pool on each H100 is managed in blocks: requests allocate on demand, and physical memory stays compact.
The story behind the name
PagedAttention is the name from the vLLM paper, but the idea of paged KV caches is not original to vLLM — the OS-virtual-memory analogy already appeared in 2022 papers. TRT-LLM and vLLM "rediscovered the same principle in parallel"; the implementation details differ, and TRT-LLM's paged KV cache is more tightly integrated with FlashAttention / FP8.
3. FlashAttention Integration
TRT-LLM enables FlashAttention 2/3 by default (FA3 on Hopper), reducing the IO complexity of attention from O(n²) to O(n²/M) (where M is the SRAM size). For FlashAttention details, see Classic Papers in Depth and GPU Architecture and Optimization.
Traditional attention: S = Q·K^T (n×n matrix written to HBM) → P = softmax(S) → O = P·V
HBM reads/writes: O(n²)
FlashAttention: split Q/K/V into tiles, complete attention inside SRAM, no intermediate matrix written
HBM reads/writes: O(n²/M)4. FP8: The H100 Killer Feature
The H100 Hopper architecture natively supports FP8 (E4M3 / E5M2), with twice the theoretical throughput of BF16. TRT-LLM supports:
- FP8 weights + FP8 activations (W8A8): extreme throughput
- FP8 weights + BF16 activations (W8A16): accuracy first
- FP8 KV cache: halves KV cache memory — a must for long-context scenarios
bash
# Quantize Llama-3 to FP8 with the NeMo framework
python examples/llama/quantize.py \
--model_dir /models/Meta-Llama-3-8B \
--dtype float16 \
--qformat fp8 \
--output_dir /models/llama3-8b-fp85. Quantization Support Matrix
| Method | Bit width | TRT-LLM support | Notes |
|---|---|---|---|
| SmoothQuant | INT8 | ✅ | Activations smoothed, W8A8, near-lossless |
| GPTQ | INT4 | ✅ | Paired with Marlin kernels |
| AWQ | INT4 | ✅ | Same as vLLM |
| FP8 | FP8 | ✅ | Recommended on H100 |
| INT4 AWQ + FP8 KV | INT4 + FP8 | ✅ | First choice for long context |
| INT8 KV cache | INT8 | ✅ | Long context on non-H100 cards |
See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
3. Build Flow: Model → Engine
TRT-LLM splits the build into two steps: first quantize / convert, then build the engine.
bash
# Step 1: convert + quantize (NeMo or TRT-LLM's native tools)
python tools/llm/llama/convert.py \
--model_dir /models/Meta-Llama-3-8B \
--output_dir /models/llama3-8b-trt-checkpoint \
--dtype float16 \
--use_smooth_quant # or --quantize_fp8 / --use_gptq / --use_awq
# Step 2: build the engine
trtllm-build \
--checkpoint_dir /models/llama3-8b-trt-checkpoint \
--output_dir /models/llama3-8b-engine \
--gemm_plugin float16 \
--gpt_attention_plugin float16 \
--context_fmha enable \
--paged_kv_cache enable \
--remove_input_padding enable \
--use_paged_context_fmha enable \
--max_batch_size 32 \
--max_input_len 2048 \
--max_output_len 512 \
--max_num_tokens 8192 \
--builder_optimization_level 3Key Build Switches
| Switch | Purpose | Recommendation |
|---|---|---|
--gemm_plugin | Use TRT-LLM's own GEMM (replacing TRT's default) | float16 / bfloat16 / fp8 |
--gpt_attention_plugin | Use the FlashAttention variant | float16 / fp8 |
--context_fmha | FlashAttention in the prefill phase | enable (always on) |
--use_paged_context_fmha | Paged KV in prefill as well | Recommended on H100 |
--paged_kv_cache | Paged KV cache | enable (always on) |
--remove_input_padding | Remove padding, packed sequences | enable (always on) |
--max_num_tokens | Max tokens per step | Determines the batch × seq ceiling |
--builder_optimization_level | Optimization depth (0–5) | 3–5, trading build time for performance |
Builds take a long time
--builder_optimization_level 5 + a large model → a single build takes hours to tens of hours. CI must cache engines (named by model hash) and never rebuild every time. Same as TensorRT.
4. Performance Tuning: Measured Paths
Here is a set of baselines on H100 80GB (see Inference Benchmarking in Practice):
| Model | Precision | Concurrency | Throughput (tokens/s) | First-token latency |
|---|---|---|---|---|
| Llama-3-8B | BF16 | 32 | ~5500 | 28 ms |
| Llama-3-8B | FP8 (W8A8) | 32 | ~10000 | 28 ms |
| Llama-3-8B | INT4 AWQ + FP8 KV | 32 | ~12000 | 30 ms |
| Llama-3-70B | BF16 (2×H100 TP) | 32 | ~1700 | 75 ms |
| Llama-3-70B | FP8 (1×H100) | 32 | ~1900 | 70 ms |
| Mixtral-8x22B | FP8 (4×H100 TP) | 32 | ~2000 | 120 ms |
Against vLLM 0.6.x (same hardware):
Llama-3-8B FP8: TRT-LLM ~10000 tokens/s vs vLLM ~8500 tokens/s (1.18×)
Llama-3-70B FP8 TP=2: TRT-LLM ~1900 tokens/s vs vLLM ~1700 tokens/s (1.12×)TRT-LLM is slightly faster on H100; on A100 / L40 and similar cards the gap narrows or even reverses. Don't assume the outcome — benchmark first.
5. Tuning in Practice: Common Switches
1. Long-Context Optimization
bash
# Llama-3.1-8B-Instruct with 128K context
trtllm-build \
... \
--max_input_len 127000 \
--max_output_len 1000 \
--max_num_tokens 131072 \
--context_fmha enable \
--use_paged_context_fmha enable \
--paged_kv_cache enable \
--kv_cache_quant_type INT8 # or FP8, halving the KV cacheThe long-context bottleneck is KV cache memory. The KV cache of Llama-3-8B at 128K context is ~40 GB, or ~20 GB after FP8 quantization — fitting on a single card.
2. Speculative Decoding
TRT-LLM 0.10+ supports speculative decoding (see Speculative Decoding and Medusa/EAGLE):
bash
# EAGLE speculative decoding
trtllm-build \
... \
--speculative_decoding_mode eagle \
--eagle_model_dir /models/llama3-8b-eagleEAGLE-3 achieves 2–4× on SGLang/vLLM; TRT-LLM is similar.
3. Triton Deployment
The official deployment form of TRT-LLM is the Triton LLM Backend (see Triton Inference Server):
models/
└── llama3-8b/
├── 1/
│ └── (path to the TRT-LLM engine)
└── config.pbtxtconfig.pbtxt:
backend: "tensorrtllm"
max_batch_size: 32
input [
{ name: "input_ids", data_type: TYPE_INT32, dims: [-1] },
{ name: "input_lengths", data_type: TYPE_INT32, dims: [1] },
{ name: "request_output_len", data_type: TYPE_INT32, dims: [-1] },
{ name: "stop_words_list", data_type: TYPE_INT32, dims: [-1] }
]
output [
{ name: "output_ids", data_type: TYPE_INT32, dims: [-1, -1] },
{ name: "output_lengths", data_type: TYPE_INT32, dims: [-1] },
{ name: "cum_log_probs", data_type: TYPE_FP32, dims: [-1] }
]
parameters: {
key: "decoder_model_path"
value: { string_value: "/models/llama3-8b/engine" }
}
parameters: {
key: "max_beam_width"
value: { string_value: "1" }
}
parameters: {
key: "TRTLLM_Batch_scheduler_policy"
value: { string_value: "max_utilization" }
}6. Comparison with vLLM
| Dimension | TensorRT-LLM | vLLM |
|---|---|---|
| H100 performance | 1.0–1.3×, slightly ahead | The baseline |
| A100 performance | Comparable | Comparable |
| Ease of use | Heavy build flow, steep learning curve | One line to start a server |
| Model support | Mainstream models covered; new models wait for NVIDIA adaptation | HuggingFace models out of the box |
| Quantization | FP8 / SmoothQuant / GPTQ / AWQ | AWQ / GPTQ / FP8 / bitsandbytes |
| Community | NVIDIA-led, enterprise-heavy | UC Berkeley-led, community-driven |
| Structured output | outlines / xgrammar integration | The SGLang family is stronger |
| Multimodal | Supported but relatively behind | Decent since 0.5+ |
| Cross-hardware | NVIDIA only | NVIDIA only (same) |
| Triton integration | Native LLM Backend | Via the vLLM Triton backend |
Decision guide
- H100 cluster, chasing the extreme: try TRT-LLM first and keep it if benchmarks favor it; if the difference is small, use vLLM (ease-of-use advantage).
- A100 / L40 / T4: vLLM is usually the less troublesome choice.
- Quick PoCs, small teams: vLLM.
- NVIDIA partners / enterprise production: the full TRT-LLM + Triton stack, backed by NVIDIA commercial support.
- Structured output / agents: consider SGLang instead of TRT-LLM.
7. Limitations and Boundaries
- Steep learning curve: dozens of build switches and documentation scattered across GitHub issues; beginners give up easily.
- Long build times: kernel auto-tuning is extremely slow; CI must cache engines.
- Lagging support for new models: new architectures (Mamba, Hyena, new MoE variants) land 1–2 months behind vLLM.
- Bound to NVIDIA: unavailable for cross-hardware scenarios — go with ONNX Runtime or llama.cpp.
- Complex Triton configuration: a miswritten config.pbtxt is an easy trap; you need to be familiar with the Triton LLM backend protocol.
- Closed-source components: the core plugins are closed, making debugging hard.
8. Where to Go Next
- Concept pages: Batching and Request Scheduling, The GPU Memory Hierarchy and the Bandwidth Wall, Model Quantization Fundamentals, Kernel Fusion and Custom Kernels, Computation Graph Optimization, GPU Architecture and Optimization, Latency, Throughput, and Concurrency, Model Serving and Orchestration
- Case-study pages: TensorRT, vLLM, Speculative Decoding and Medusa/EAGLE, Triton Inference Server, Distributed Inference (TP/PP)
- Papers: Classic Papers in Depth, Reading Paths, Frontier Advances
- Practice pages: Inference Engine Comparison, Tuning and Performance Optimization, Inference Benchmarking in Practice, Common Pitfalls and Anti-Patterns
- Resource pages: Hardware Primer, Benchmark Data & Tool Profiles, Curated Resources
References
- NVIDIA. TensorRT-LLM official documentation — entry point
- NVIDIA. TensorRT-LLM GitHub — source code and examples
- NVIDIA. Deploying LLMs with TensorRT-LLM, Triton and vLLM (GTC 2024) — official comparison
- Dao et al. FlashAttention-2 (2023) — the FA2 paper
- Xiao et al. SmoothQuant: Accurate and Efficient Post-Training Quantization for LLMs (ICML 2023) — SmoothQuant
- Xiao et al. Efficient Generative LLM Inference Using Phase Prediction (ISCA 2024) — the In-Flight Batching engineering paper
- Cai et al. Medusa (ICML 2024) — speculative decoding also supported by TRT-LLM
- Li et al. EAGLE-3 (2025) — the latest speculative decoding