Skip to content

Triton Inference Server

At a glance NVIDIA's open-source multi-framework, multi-model, multi-backend inference server — the de facto standard for taking LLMs and traditional models to production. This article dissects the Backend concept, config.pbtxt, concurrent execution, dynamic batching, the model repository, Prometheus metrics, and KServe integration.

Triton Inference Server ​

1. Definition: Turning an "Inference Engine" into an "Inference Service" ​

Triton Inference Server is NVIDIA's open-source multi-framework, multi-model, multi-backend inference server. Its positioning: wrap a single-machine inference engine (TensorRT / ONNX Runtime / PyTorch / TensorFlow / custom Python) into a horizontally scalable HTTP/gRPC service — you give it a model repository, and it gives you a production-ready inference API.

Understand Triton's place in the inference stack:

                ┌──────────────────────────────────────────┐
                │   Clients (Web / App / backend services) │
                └────────────────┬─────────────────────────┘
                                 │ HTTP / gRPC (KServe v2)
                                 ▼
                ┌──────────────────────────────────────────┐
                │  Triton Inference Server                 │
                │  ├─ Dynamic batching                     │
                │  ├─ Concurrent model execution           │
                │  ├─ Model repository + versioning        │
                │  ├─ Metrics (Prometheus)                 │
                │  └─ Backend abstraction (TRT/ORT/PyTorch)│
                └────────────────┬─────────────────────────┘
                                 │ internal calls
                                 ▼
        ┌────────────────────────────────────────────────────┐
        │  Backends                                           │
        │  [TensorRT] [ONNX Runtime] [PyTorch] [TF] [Python]  │
        │  [Fil] [vLLM] [TensorRT-LLM] [OpenVINO] ...         │
        └─────────────────────────────────────────────────────┘

Triton is the "serving shell" for TensorRT, ONNX Runtime, TensorRT-LLM, and vLLM — each engine has its strengths, but none can simply "open a port and go live." What Triton adds is the protocol layer, the batching layer, the model-management layer, and the observability layer. See Model Serving and Orchestration and Batching and Request Scheduling.

2. Backend: The Pluggable Engine Abstraction ​

Triton's core abstraction is the Backend: each backend is an "inference engine adapter" implemented in C++ or Python. Under one model repository, different models can use different backends:

BackendPurposeNotes
tensorrtLoads TensorRT enginesSee the TensorRT case study
onnxruntimeLoads ONNX modelsSee the ONNX Runtime case study
pytorchLoads TorchScript modelsUses libtorch, no Python dependency
pythonArbitrary Python inference logicVia the Triton Python Backend, subject to the GIL
tensorflowLoads SavedModelsTF1 / TF2
filTree models (XGBoost / LightGBM / RAPIDS)GBDT deployment
openvinoIntel CPU/iGPUSee the OpenVINO case study
tensorrtllmTensorRT-LLM LLM enginesSee the TensorRT-LLM case study
vllmvLLM LLM enginesSee the vLLM case study
python (decoupled)High-performance Python backendActor model, bypassing the GIL

Each backend receives Triton's "create model instance" request and returns an instance object that can run enqueue. Multiple models can reside in parallel — this is Triton's killer feature: one serving process can simultaneously run a vision model (TensorRT), an LLM (TensorRT-LLM), and a GBDT (fil), all sharing the GPU.

3. The Model Repository and config.pbtxt ​

Triton uses the filesystem as its model repository — a directory tree:

model_repository/
├── resnet50/
│   ├── config.pbtxt              ← model configuration
│   └── 1/                        ← version 1
│       └── model.plan            ← TensorRT engine
├── bert_qa/
│   ├── config.pbtxt
│   ├── 1/
│   │   └── model.onnx            ← ONNX model
│   └── 2/
│       └── model.onnx            ← new version
├── llama3-8b/
│   ├── config.pbtxt
│   └── 1/
│       └── (path to the TRT-LLM engine)
└── gbdt_ranker/
    ├── config.pbtxt
    └── 1/
        └── model.json            ← XGBoost model

A config.pbtxt Example ​

name: "resnet50"
backend: "tensorrt"
max_batch_size: 128

input [
  {
    name: "input"
    data_type: TYPE_FP32
    dims: [ 3, 224, 224 ]
  }
]
output [
  {
    name: "output"
    data_type: TYPE_FP32
    dims: [ 1000 ]
  }
]

dynamic_batching {
  preferred_batch_size: [ 4, 8, 16, 32, 64 ]
  max_queue_delay_microseconds: 50000      # must form a batch within 50ms
  preserve_ordering: true
}

instance_group [
  {
    kind: KIND_GPU
    count: 2                                # 2 concurrent instances
    gpus: [ 0 ]
  }
]

parameters: {
  key: "workspace_size"
  value: { string_value: "1073741824" }    # 1 GB workspace
}

Key fields:

  • max_batch_size: maximum batch, setting the per-request batch ceiling
  • dynamic_batching: dynamic batching configuration (see the next section)
  • instance_group: number of model instances and GPU assignment
  • parameters: backend-specific parameters

Version management

Triton uses the directories 1/, 2/ for model versions, controllable via --model-version-policy=all / specific: 1,2 / latest: 2. Production practices:

  1. Put the new version in 2/ first and set version_policy: latest:2 in config.pbtxt;
  2. Run a canary (10% of traffic) by having clients select the version number;
  3. Delete 1/ only after the full cutover;
  4. Always keep the previous version for rollback.

4. Concurrent Model Execution: Multiple Instances in Parallel ​

Concurrent Model Execution lets the same model reside as multiple instances running in parallel:

Request A ──┐
            ├──→ [Instance 1 on GPU 0] → Response A
Request B ──┤
            ├──→ [Instance 2 on GPU 0] → Response B
Request C ──┘
  • kind: KIND_GPU: runs on the GPU
  • count: N: N instances (subject to GPU SM resource constraints)
  • gpus: [0, 1]: which GPUs to place them on

More instances → higher concurrent throughput, but each instance gets fewer SMs → single-request latency may rise. Instance count is the trade-off between single-request latency and concurrent throughput — see Latency, Throughput, and Concurrency.

In LLM scenarios, the TensorRT-LLM / vLLM backends already implement PagedAttention + continuous batching internally, so a single instance delivers extremely high throughput — extra instances are usually unnecessary (they just fight for resources).

5. Dynamic Batching: Pooling Scattered Requests into Batches ​

Dynamic batching is Triton's signature feature: scattered client requests "wait for a short while" on the server side to form a batch, then run in one forward pass:

T=0ms:   Request A (batch=1)
T=2ms:   Request B (batch=1)
T=5ms:   Request C (batch=1)
T=10ms:  3 requests pooled → one forward pass with batch=3 → responses A,B,C
T=50ms:  even if the batch isn't full, force it to run (max_queue_delay)

Key parameters:

  • preferred_batch_size: [4, 8, 16, 32, 64]: run immediately once the batch reaches these sizes
  • max_queue_delay_microseconds: 50000: wait at most 50 ms
  • preserve_ordering: true: keep responses in request order (recommended for LLMs)

The dynamic batching trade-off

  • Bigger batches → higher throughput, but higher latency (requests wait to be pooled);
  • Smaller max_queue_delay → lower latency, but smaller batches;
  • Vision / NLP classification: max_queue_delay = 50 ms, preferred = [8, 16, 32] is usually a good starting point;
  • LLMs: dynamic batching usually doesn't apply, because continuous batching (see vLLM / TensorRT-LLM) already pools at iteration level.

6. Protocols and Clients ​

Triton supports two protocols:

  • HTTP / REST: JSON, easy to debug
  • gRPC: protobuf, higher performance, recommended for production
python
# Python client
import tritonclient.http as httpclient

client = httpclient.InferenceServerClient(url="localhost:8000")

# Check model readiness
assert client.is_model_ready("resnet50")

# Inference
inputs = [httpclient.InferInput("input", [1, 3, 224, 224], "FP32")]
inputs[0].set_data_from_numpy(input_array)

outputs = [httpclient.InferRequestedOutput("output")]

response = client.infer("resnet50", inputs, outputs=outputs)
result = response.as_numpy("output")

Triton is compatible with the KServe v2 protocol (see Model Serving and Orchestration), so it can be called directly by upper-layer frameworks such as KServe / Seldon Core / BentoML.

7. Metrics and Observability ​

Triton ships with a Prometheus-format metrics endpoint:

GET http://localhost:8002/metrics

Key metrics:

MetricMeaning
nv_inference_request_successSuccessful request count
nv_inference_request_failureFailed request count
nv_inference_request_durationTotal request time (including queuing)
nv_inference_infer_durationActual inference time
nv_inference_queue_durationQueue wait time
nv_inference_exec_countModel execution count
nv_gpu_memory_used_bytesGPU memory usage

Triton provides ready-made Grafana dashboards (see Curated Resources and Common Pitfalls and Anti-Patterns).

Metrics you must watch before going live

  1. queue_duration 90th percentile > 100 ms → batching too aggressive, lower preferred_batch_size;
  2. infer_duration / request_duration ratio < 0.5 → batching/serialization overhead too high, switch to gRPC;
  3. Rising failures → mostly OOM or timeouts; check GPU memory;
  4. Severely uneven exec_count across model instances → uneven traffic; consider more instances or a load balancer.

8. LLM Deployment: Triton + TensorRT-LLM / vLLM ​

LLM inference differs from small-to-medium models — PagedAttention, continuous batching, and KV cache management require dedicated backends. Triton supports LLMs through the tensorrtllm backend and the vllm backend:

TRT-LLM Deployment (see TensorRT-LLM) ​

models/
└── llama3-8b/
    ├── 1/
    │   ├── config.json
    │   └── engine/                     ← TRT-LLM engine
    └── config.pbtxt

config.pbtxt:

backend: "tensorrtllm"
max_batch_size: 0                        # the LLM manages its own batching; Triton should not batch

input [
  { name: "text_input",      data_type: TYPE_STRING, dims: [1] },
  { name: "max_tokens",      data_type: TYPE_INT32,  dims: [1] },
  { name: "temperature",     data_type: TYPE_FP32,  dims: [1] }
]
output [
  { name: "text_output",     data_type: TYPE_STRING, dims: [1] }
]

parameters: {
  key: "TRTLLM_Batch_scheduler_policy"
  value: { string_value: "max_utilization" }
}

Launch:

bash
tritonserver \
    --model-repository=/models \
    --http-port=8000 \
    --grpc-port=8001 \
    --metrics-port=8002

vLLM Backend ​

Since 2024, Triton also supports the vLLM backend. Deployment is similar but simpler — it can consume HuggingFace model IDs directly.

9. Integration with KServe / Seldon Core ​

Triton is a "single-process inference server"; KServe / Seldon Core are "inference platforms on Kubernetes." The two work together:

┌────────────────────────────────────────────────┐
│  Kubernetes                                    │
│                                                │
│  ┌──────────────────────────────────────────┐  │
│  │  KServe InferenceService                │  │
│  │  ├─ Knative Service (auto scaling)       │  │
│  │  ├─ Routing, ingress, TLS                │  │
│  │  └─ Container                            │  │
│  │     └─ tritonserver (the topic here)     │  │
│  └──────────────────────────────────────────┘  │
│                                                │
└────────────────────────────────────────────────┘
  • KServe: handles deployment on K8s, autoscaling, traffic management, and canary rollouts
  • Triton: handles inference, batching, and multi-model management within a single pod

See Model Serving and Orchestration and Portfolio Projects.

10. Performance Data: Baseline Reference ​

Triton + TensorRT-LLM running Llama-3-8B FP8 on a single A100 80GB:

ConcurrencyThroughput (tokens/s)First-token latencySteady-state latency (per token)
115025 ms6.7 ms
8110030 ms7.3 ms
32450045 ms7.1 ms
128850090 ms15 ms
2569500180 ms27 ms

Observations:

  • At concurrency 32, throughput is 4500 tokens/s with a 45 ms first token — production-ready;
  • At concurrency 256, throughput plateaus (approaching the memory-bandwidth bottleneck) and latency climbs — this is the bandwidth wall.

11. Limitations and Boundaries ​

  1. NVIDIA-centric: although it supports ONNX Runtime / OpenVINO / TF backends, the optimization depth trails the NVIDIA stack; non-NVIDIA GPUs get a poor experience.
  2. Slow Python Backend: the Python backend is limited by the GIL and performs poorly — for heavy loads, use the Decoupled Python Backend or a C++ backend.
  3. Complex configuration: config.pbtxt has dozens of fields with scattered documentation; beginners trip often.
  4. LLM model repository management: LLM engine paths and versioning are more complex than single-file models and need CI support.
  5. OOM when multiple models share a GPU: misconfigured instance counts cause OOM; test and monitor.
  6. The open-source version lacks enterprise features: things like multi-replica shared model repositories, A/B testing, and automatic canaries require upper-layer frameworks such as KServe.

12. Comparison with Similar Solutions ​

SolutionRelationship with Triton
TorchServeOfficial PyTorch offering with deep PyTorch integration; few backends, weaker performance
TF ServingOfficial TF offering with deep TF integration; similar limits to TorchServe
BentoMLPython-first with a great developer experience; few backends, enterprise features self-assembled
KServeNot a Triton competitor but the upper-layer platform; Triton is one of KServe's runtimes
Seldon CoreSame as KServe — an upper-layer platform
OpenVINO Model ServerSee OpenVINO; Intel-centric but functionally similar
vLLM servervLLM ships its own OpenAI-compatible server — lighter than Triton for single-LLM scenarios; use Triton for multi-model mixes
Candle / candle-coreRust ecosystem; clean but small

13. Where to Go Next ​

References ​