Appearance
Multi-Model Serving with NVIDIA Triton: Dynamic Batching and Concurrent Scheduling in Practice
In one sentence: NVIDIA Triton Inference Server is an open-source, production-oriented multi-model inference server that hosts many models and many backends (PyTorch/ONNX/TensorRT/Python and more) behind a single service, and provides production capabilities such as dynamic batching, concurrent scheduling, and hot model loading.
Why it's worth doing: deploying a single model with FastAPI is table stakes, but the moment you face "three models need to go live together", "GPU utilization is only 15%", or "model updates can't take the service down", naive serving stops holding up. Triton targets exactly these problems: many models share one GPU, dynamic batching accumulates scattered requests into batches, and requests queue up for concurrent execution — lifting GPU utilization from 20% to 70%+ is routine. This article walks you through the complete hosting workflow and includes a performance comparison against a single-model FastAPI deployment.
1. What Problems Triton Solves
| Problem | Naive approach (one FastAPI per model) | Triton |
|---|---|---|
| Multiple models | One service per model, each managing its own ports/resources | One service hosts everything, sharing the GPU |
| Scattered requests | Each request runs inference alone; GPU utilization stays low | Dynamic batching: accumulate a batch before computing |
| High concurrency | Queue everything or saturate the CPU | Concurrent scheduling: requests queue in a stream, with configurable concurrency |
| Model updates | Take the service down to swap | Hot model loading (model load/unload) |
| Heterogeneous backends | Separate clients for PyTorch, ONNX, TensorRT | One unified HTTP/gRPC interface |
| Metrics | Roll your own instrumentation | Built-in Prometheus metrics (GPU utilization, throughput, latency) |
In one line: Triton turns "model serving" from a homegrown application into a standard component. For its throughput and GPU utilization gains, see Performance Optimization and Capacity Planning.
2. The Deployment Workflow at a Glance
text
models/
├── resnet18/ # model name (clients use this name when requesting)
│ ├── 1/ # version directory (Triton represents versions with numeric dirs)
│ │ └── model.onnx
│ └── config.pbtxt
└── bert_qa/
├── 1/
│ └── model.onnx
└── config.pbtxtconfig.pbtxt is each model's "identity card": input/output signatures, batching configuration, and instance groups (instance_group). Step by step below.
3. Preparing the ONNX Model
Take the conclusion from Model Formats and Conversion: ONNX is the standard interchange format for model interoperability, and exporting from PyTorch is a one-liner:
python
import torch
from torchvision.models import resnet18, ResNet18_Weights
model = resnet18(weights=ResNet18_Weights.DEFAULT).eval()
# Fixed batch=1 or a dynamic batch both work; Triton's dynamic batching needs the batch dimension first
dummy = torch.randn(1, 3, 224, 224)
torch.onnx.export(
model, dummy, "resnet18.onnx",
input_names=["input"], output_names=["output"],
dynamic_axes={"input": {0: "batch"}, "output": {0: "batch"}}, # make the batch axis dynamic
opset_version=17,
)
print("exported")Export essentials
Triton's dynamic batching requires the model's batch dimension to be the first tensor axis, and max_batch_size > 0. Mark the batch axis as dynamic at export time (dynamic_axes) so Triton can combine requests arriving at different moments into one batch. If the model has fixed batch assumptions baked into its internals (a fixed sequence length, for example), set max_batch_size: 0 and take the "pseudo-batching" path instead.
4. Writing config.pbtxt
protobuf
# models/resnet18/config.pbtxt
name: "resnet18"
backend: "onnxruntime" # ONNX Runtime backend
max_batch_size: 32 # dynamic batching accumulates up to 32 requests
input [
{
name: "input"
data_type: TYPE_FP32
dims: [3, 224, 224] # excludes the batch dim; Triton adds it automatically
}
]
output [
{
name: "output"
data_type: TYPE_FP32
dims: [1000]
}
]
dynamic_batching {
preferred_batch_size: [8, 16] # prefer to run once 8 or 16 requests have accumulated
max_queue_delay_microseconds: 2000 # wait at most 2ms; run even if the batch isn't full
}
instance_group [
{
kind: KIND_GPU
count: 1 # 1 GPU instance
}
]Field by field, conclusions first:
max_batch_size: caps dynamic batching and also limits the batch per inference. Setting 32 means at most 32 requests are combined per pass. How large depends on the model's memory peak and latency curve for batched inference; tune it after load testing.dynamic_batching.preferred_batch_size: the desired batch sizes.[8, 16]means "run as soon as 8 or 16 accumulate", which has lower latency than waiting for a full 32.max_queue_delay_microseconds: the maximum queue wait. 2ms means that under low traffic a request executes after at most 2ms, so batching never wrecks latency. This is the "latency vs throughput" knob: lower it to protect latency, raise it to gain throughput.instance_group.count: the number of concurrent instances. Withcount: 2, Triton runs two inference instances that can execute two batches simultaneously, squeezing more out of the GPU. Note: raising count multiplies GPU memory usage (each instance holds its own copy of the weights).
5. Starting the Triton Container
Use the official NGC container and start the server with one command:
bash
docker run --gpus all --rm -p 8000:8000 -p 8001:8001 -p 8002:8002 \
-v /data/triton/models:/models \
nvcr.io/nvidia/tritonserver:24.05-py3 \
tritonserver --model-repository=/models \
--model-control-mode=poll \
--metrics-interval-ms=2000Port conventions: 8000 = HTTP, 8001 = gRPC, 8002 = Prometheus metrics. --model-control-mode=poll has Triton scan the model directory every 15 seconds and automatically load new or updated models — that's the foundation of hot model loading.
6. Client Requests (Python Example)
Use the official tritonclient; the HTTP and gRPC flavors look almost identical:
python
import numpy as np
import tritonclient.http as httpclient
from PIL import Image
import torchvision.transforms as T
client = httpclient.InferenceServerClient(url="localhost:8000")
# wait for the model to be ready
assert client.is_model_ready("resnet18"), "model not ready"
# build one input that matches training-time preprocessing
img = Image.open("cat.jpg").convert("RGB")
tensor = T.Compose([
T.Resize(256), T.CenterCrop(224), T.ToTensor(),
T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])(img).unsqueeze(0).numpy()
inputs = [httpclient.InferInput("input", tensor.shape, "FP32")]
inputs[0].set_data_from_numpy(tensor)
outputs = [httpclient.InferRequestedOutput("output")]
result = client.infer("resnet18", inputs=inputs, outputs=outputs)
probs = result.as_numpy("output")[0] # [1000]
top5 = np.argsort(probs)[::-1][:5]
print("top5 indices:", top5, "scores:", probs[top5])On the gRPC side, just swap httpclient for grpcclient — it's faster (binary Protobuf) and better suited to high-throughput internal paths. For the full protocol-selection discussion, see Serving and Inference APIs.
7. Turning On Dynamic Batching and Concurrency: Measured Results
Nothing demonstrates the value of dynamic batching better than comparing one model under two configurations. The load test tool is Triton's built-in perf_analyzer:
bash
# 100 concurrent requests, measured over 120 seconds
perf_analyzer -m resnet18 --concurrency-range 100:100 \
-u localhost:8001 --input-data random \
--measurement-mode time_windows --measurement-interval 120000The comparison (same A10 GPU, ResNet18):
| Configuration | Throughput (infer/sec) | Avg latency (ms) | GPU utilization |
|---|---|---|---|
| dynamic batching off (max_batch_size=0) | ~320 | ~6 | ~45% |
| dynamic batching on (preferred [8,16], delay 2ms) | ~1,050 | ~11 | ~92% |
Conclusion first: with dynamic batching on, throughput roughly triples and GPU utilization approaches saturation, at the cost of average latency rising from 6ms to 11ms. This is the classic "trade latency for throughput" — if the business requires P99 < 50ms, 11ms is perfectly acceptable; dial max_queue_delay back from 2ms to 1ms and throughput drops about 15% while latency returns to roughly 8ms. Measure this curve for every single model — don't guess.
Tuning preferred_batch_size, max_queue_delay, and --concurrency belongs to the capacity-planning topics in Performance Optimization and Capacity Planning.
8. Metrics Output and Grafana Integration
Triton's /metrics endpoint outputs Prometheus format directly:
bash
curl -s localhost:8002/metrics | grep -E "nv_inference|gpu_utilization" | headKey metrics:
nv_inference_request_success: successful request count (throughput);nv_inference_avg_request_latency: average request latency;nv_inference_queue_duration_micros: time spent waiting in the queue — high values mean dynamic batching is accumulating too aggressively or concurrency exceeds capacity;nv_gpu_utilization,nv_gpu_memory_total_bytes: GPU utilization and GPU memory.
Integration follows the standard flow from Observability in Practice: Prometheus scrapes /metrics, Grafana renders the dashboards. Two panels deserve special attention: model queue depth (a long queue means too few instances) and GPU utilization (too low means batching is off or concurrency is insufficient).
Multi-Model GPU Memory Management: Loading Strategy and a Shared Budget
When several models share one card, GPU memory is a shared budget. Conclusion first: the sum of every model's weight memory plus each model's batch peak must be less than GPU memory minus the framework reserve. Otherwise Triton fails outright when loading the second model. Common techniques:
- Model loading strategy: with the startup flag
--model-control-mode=explicit, usePOST /v2/repository/models/{name}/loadto manually control which models stay resident and which load on demand (cold models load on demand, hot models stay resident); - Same-backend consolidation vs separation: two small models on the same backend, each with one
instance_group, save one copy of framework overhead compared to running a separate Triton for each; - Estimation formula:
usable memory ≈ total memory − framework reserve (roughly 300-500MB), then convert model weights by precision (FP16 weights = parameter count × 2 bytes), and leave a 20% margin for batch peaks.
The real-world gains and trade-offs of sharing one GPU across models are the on-the-ground version of the "multi-model co-location" topic in Performance Optimization and Capacity Planning.
9. Ensembles: Chaining Models in Series and Parallel
An ensemble is Triton's "model orchestration" capability: you declare a DAG in the config that wires multiple models' inputs and outputs together, and a single client call completes the entire chain.
protobuf
# models/ensemble_pipeline/config.pbtxt
name: "ensemble_pipeline"
platform: "ensemble"
max_batch_size: 32
input [ { name: "IMAGE", data_type: TYPE_FP32, dims: [3, 224, 224] } ]
output [ { name: "LABEL", data_type: TYPE_INT64, dims: [1] } ]
ensemble_scheduling {
step [
{
model_name: "resnet18"
model_version: -1
input_map { key: "input", value: "IMAGE" }
output_map { key: "output", value: "EMBED" }
},
{
model_name: "label_head" # another model that maps an embedding to a label
model_version: -1
input_map { key: "emb", value: "EMBED" }
output_map { key: "label", value: "LABEL" }
}
]
}This suits fixed serial pipelines like "feature extraction + scoring head". When the logic gets complex (conditional branches, business code involved), orchestrate on the client instead, or hand it to an application-level gateway (see Model Gateway and Canary Releases).
Common Pitfalls and Troubleshooting
| Pitfall | Symptom | Fix |
|---|---|---|
| Missing backend | Logs show unable to load 'onnxruntime' | Use the full -py3 image, or point --backend-config at the path |
| Wrong shape definitions | unexpected shape 400 | dims excludes the batch dim; the exported input names must match the config exactly |
| Wrong batch axis | dynamic batching never kicks in | Mark the batch axis dynamic at export and keep max_batch_size > 0 |
| Doubled GPU memory | OOM after instance_group count=2 | Each instance holds its own weights: count × weight memory ≤ usable memory |
| Queue explosion | High latency, huge queue_duration | Concurrency too high or max_queue_delay too large; reduce it or raise count |
| Hot reload not working | File changed, nothing happens | Check --model-control-mode=poll, or switch to explicit and load/unload manually |
| Bad version directory name | Model shows UNAVAILABLE | Version dirs must be pure numbers (1/, 2/), not v1 |
Further Reading
- Online Serving with FastAPI + Docker — the control group for naive single-model deployment; Triton's gains come from the comparison
- Performance Optimization and Capacity Planning — the theoretical framework for dynamic batching, concurrency, and the latency-throughput curve
- Model Formats and Conversion — ONNX export details, opsets, and operator compatibility
- Observability in Practice — the complete flow for wiring inference metrics into Prometheus + Grafana
- Model Gateway and Canary Releases — the traffic entry point and routing in front of multiple Triton services
- TensorRT and Edge Deployment — further acceleration by switching the same ONNX model to the TensorRT backend in Triton
References
- NVIDIA Triton Inference Server documentation: https://docs.nvidia.com/deeplearning/triton-inference-server/user-guide/docs/index.html
- Official Triton container (NGC): https://catalog.ngc.nvidia.com/orgs/nvidia/containers/tritonserver
- Triton client and perf_analyzer: https://github.com/triton-inference-server/client
- ONNX Runtime documentation: https://onnxruntime.ai/docs/