Skip to content

Multi-Model Serving with NVIDIA Triton

At a glance A complete hands-on guide to hosting multiple models on NVIDIA Triton Inference Server with dynamic batching and concurrent scheduling enabled: model repository layout, writing config.pbtxt, performance metrics, and load-test comparisons.

Multi-Model Serving with NVIDIA Triton: Dynamic Batching and Concurrent Scheduling in Practice ​

In one sentence: NVIDIA Triton Inference Server is an open-source, production-oriented multi-model inference server that hosts many models and many backends (PyTorch/ONNX/TensorRT/Python and more) behind a single service, and provides production capabilities such as dynamic batching, concurrent scheduling, and hot model loading.

Why it's worth doing: deploying a single model with FastAPI is table stakes, but the moment you face "three models need to go live together", "GPU utilization is only 15%", or "model updates can't take the service down", naive serving stops holding up. Triton targets exactly these problems: many models share one GPU, dynamic batching accumulates scattered requests into batches, and requests queue up for concurrent execution — lifting GPU utilization from 20% to 70%+ is routine. This article walks you through the complete hosting workflow and includes a performance comparison against a single-model FastAPI deployment.

1. What Problems Triton Solves ​

ProblemNaive approach (one FastAPI per model)Triton
Multiple modelsOne service per model, each managing its own ports/resourcesOne service hosts everything, sharing the GPU
Scattered requestsEach request runs inference alone; GPU utilization stays lowDynamic batching: accumulate a batch before computing
High concurrencyQueue everything or saturate the CPUConcurrent scheduling: requests queue in a stream, with configurable concurrency
Model updatesTake the service down to swapHot model loading (model load/unload)
Heterogeneous backendsSeparate clients for PyTorch, ONNX, TensorRTOne unified HTTP/gRPC interface
MetricsRoll your own instrumentationBuilt-in Prometheus metrics (GPU utilization, throughput, latency)

In one line: Triton turns "model serving" from a homegrown application into a standard component. For its throughput and GPU utilization gains, see Performance Optimization and Capacity Planning.

2. The Deployment Workflow at a Glance ​

text
models/
├── resnet18/                 # model name (clients use this name when requesting)
│   ├── 1/                    # version directory (Triton represents versions with numeric dirs)
│   │   └── model.onnx
│   └── config.pbtxt
└── bert_qa/
    ├── 1/
    │   └── model.onnx
    └── config.pbtxt

config.pbtxt is each model's "identity card": input/output signatures, batching configuration, and instance groups (instance_group). Step by step below.

3. Preparing the ONNX Model ​

Take the conclusion from Model Formats and Conversion: ONNX is the standard interchange format for model interoperability, and exporting from PyTorch is a one-liner:

python
import torch
from torchvision.models import resnet18, ResNet18_Weights

model = resnet18(weights=ResNet18_Weights.DEFAULT).eval()

# Fixed batch=1 or a dynamic batch both work; Triton's dynamic batching needs the batch dimension first
dummy = torch.randn(1, 3, 224, 224)
torch.onnx.export(
    model, dummy, "resnet18.onnx",
    input_names=["input"], output_names=["output"],
    dynamic_axes={"input": {0: "batch"}, "output": {0: "batch"}},  # make the batch axis dynamic
    opset_version=17,
)
print("exported")

Export essentials

Triton's dynamic batching requires the model's batch dimension to be the first tensor axis, and max_batch_size > 0. Mark the batch axis as dynamic at export time (dynamic_axes) so Triton can combine requests arriving at different moments into one batch. If the model has fixed batch assumptions baked into its internals (a fixed sequence length, for example), set max_batch_size: 0 and take the "pseudo-batching" path instead.

4. Writing config.pbtxt ​

protobuf
# models/resnet18/config.pbtxt
name: "resnet18"
backend: "onnxruntime"          # ONNX Runtime backend
max_batch_size: 32              # dynamic batching accumulates up to 32 requests

input [
  {
    name: "input"
    data_type: TYPE_FP32
    dims: [3, 224, 224]         # excludes the batch dim; Triton adds it automatically
  }
]
output [
  {
    name: "output"
    data_type: TYPE_FP32
    dims: [1000]
  }
]

dynamic_batching {
  preferred_batch_size: [8, 16]   # prefer to run once 8 or 16 requests have accumulated
  max_queue_delay_microseconds: 2000  # wait at most 2ms; run even if the batch isn't full
}

instance_group [
  {
    kind: KIND_GPU
    count: 1                       # 1 GPU instance
  }
]

Field by field, conclusions first:

  • max_batch_size: caps dynamic batching and also limits the batch per inference. Setting 32 means at most 32 requests are combined per pass. How large depends on the model's memory peak and latency curve for batched inference; tune it after load testing.
  • dynamic_batching.preferred_batch_size: the desired batch sizes. [8, 16] means "run as soon as 8 or 16 accumulate", which has lower latency than waiting for a full 32.
  • max_queue_delay_microseconds: the maximum queue wait. 2ms means that under low traffic a request executes after at most 2ms, so batching never wrecks latency. This is the "latency vs throughput" knob: lower it to protect latency, raise it to gain throughput.
  • instance_group.count: the number of concurrent instances. With count: 2, Triton runs two inference instances that can execute two batches simultaneously, squeezing more out of the GPU. Note: raising count multiplies GPU memory usage (each instance holds its own copy of the weights).

5. Starting the Triton Container ​

Use the official NGC container and start the server with one command:

bash
docker run --gpus all --rm -p 8000:8000 -p 8001:8001 -p 8002:8002 \
  -v /data/triton/models:/models \
  nvcr.io/nvidia/tritonserver:24.05-py3 \
  tritonserver --model-repository=/models \
               --model-control-mode=poll \
               --metrics-interval-ms=2000

Port conventions: 8000 = HTTP, 8001 = gRPC, 8002 = Prometheus metrics. --model-control-mode=poll has Triton scan the model directory every 15 seconds and automatically load new or updated models — that's the foundation of hot model loading.

6. Client Requests (Python Example) ​

Use the official tritonclient; the HTTP and gRPC flavors look almost identical:

python
import numpy as np
import tritonclient.http as httpclient
from PIL import Image
import torchvision.transforms as T

client = httpclient.InferenceServerClient(url="localhost:8000")

# wait for the model to be ready
assert client.is_model_ready("resnet18"), "model not ready"

# build one input that matches training-time preprocessing
img = Image.open("cat.jpg").convert("RGB")
tensor = T.Compose([
    T.Resize(256), T.CenterCrop(224), T.ToTensor(),
    T.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])(img).unsqueeze(0).numpy()

inputs = [httpclient.InferInput("input", tensor.shape, "FP32")]
inputs[0].set_data_from_numpy(tensor)

outputs = [httpclient.InferRequestedOutput("output")]

result = client.infer("resnet18", inputs=inputs, outputs=outputs)
probs = result.as_numpy("output")[0]          # [1000]
top5 = np.argsort(probs)[::-1][:5]
print("top5 indices:", top5, "scores:", probs[top5])

On the gRPC side, just swap httpclient for grpcclient — it's faster (binary Protobuf) and better suited to high-throughput internal paths. For the full protocol-selection discussion, see Serving and Inference APIs.

7. Turning On Dynamic Batching and Concurrency: Measured Results ​

Nothing demonstrates the value of dynamic batching better than comparing one model under two configurations. The load test tool is Triton's built-in perf_analyzer:

bash
# 100 concurrent requests, measured over 120 seconds
perf_analyzer -m resnet18 --concurrency-range 100:100 \
  -u localhost:8001 --input-data random \
  --measurement-mode time_windows --measurement-interval 120000

The comparison (same A10 GPU, ResNet18):

ConfigurationThroughput (infer/sec)Avg latency (ms)GPU utilization
dynamic batching off (max_batch_size=0)~320~6~45%
dynamic batching on (preferred [8,16], delay 2ms)~1,050~11~92%

Conclusion first: with dynamic batching on, throughput roughly triples and GPU utilization approaches saturation, at the cost of average latency rising from 6ms to 11ms. This is the classic "trade latency for throughput" — if the business requires P99 < 50ms, 11ms is perfectly acceptable; dial max_queue_delay back from 2ms to 1ms and throughput drops about 15% while latency returns to roughly 8ms. Measure this curve for every single model — don't guess.

Tuning preferred_batch_size, max_queue_delay, and --concurrency belongs to the capacity-planning topics in Performance Optimization and Capacity Planning.

8. Metrics Output and Grafana Integration ​

Triton's /metrics endpoint outputs Prometheus format directly:

bash
curl -s localhost:8002/metrics | grep -E "nv_inference|gpu_utilization" | head

Key metrics:

  • nv_inference_request_success: successful request count (throughput);
  • nv_inference_avg_request_latency: average request latency;
  • nv_inference_queue_duration_micros: time spent waiting in the queue — high values mean dynamic batching is accumulating too aggressively or concurrency exceeds capacity;
  • nv_gpu_utilization, nv_gpu_memory_total_bytes: GPU utilization and GPU memory.

Integration follows the standard flow from Observability in Practice: Prometheus scrapes /metrics, Grafana renders the dashboards. Two panels deserve special attention: model queue depth (a long queue means too few instances) and GPU utilization (too low means batching is off or concurrency is insufficient).

Multi-Model GPU Memory Management: Loading Strategy and a Shared Budget ​

When several models share one card, GPU memory is a shared budget. Conclusion first: the sum of every model's weight memory plus each model's batch peak must be less than GPU memory minus the framework reserve. Otherwise Triton fails outright when loading the second model. Common techniques:

  1. Model loading strategy: with the startup flag --model-control-mode=explicit, use POST /v2/repository/models/{name}/load to manually control which models stay resident and which load on demand (cold models load on demand, hot models stay resident);
  2. Same-backend consolidation vs separation: two small models on the same backend, each with one instance_group, save one copy of framework overhead compared to running a separate Triton for each;
  3. Estimation formula: usable memory ≈ total memory − framework reserve (roughly 300-500MB), then convert model weights by precision (FP16 weights = parameter count × 2 bytes), and leave a 20% margin for batch peaks.

The real-world gains and trade-offs of sharing one GPU across models are the on-the-ground version of the "multi-model co-location" topic in Performance Optimization and Capacity Planning.

9. Ensembles: Chaining Models in Series and Parallel ​

An ensemble is Triton's "model orchestration" capability: you declare a DAG in the config that wires multiple models' inputs and outputs together, and a single client call completes the entire chain.

protobuf
# models/ensemble_pipeline/config.pbtxt
name: "ensemble_pipeline"
platform: "ensemble"
max_batch_size: 32

input [ { name: "IMAGE", data_type: TYPE_FP32, dims: [3, 224, 224] } ]
output [ { name: "LABEL", data_type: TYPE_INT64, dims: [1] } ]

ensemble_scheduling {
  step [
    {
      model_name: "resnet18"
      model_version: -1
      input_map { key: "input", value: "IMAGE" }
      output_map { key: "output", value: "EMBED" }
    },
    {
      model_name: "label_head"      # another model that maps an embedding to a label
      model_version: -1
      input_map { key: "emb", value: "EMBED" }
      output_map { key: "label", value: "LABEL" }
    }
  ]
}

This suits fixed serial pipelines like "feature extraction + scoring head". When the logic gets complex (conditional branches, business code involved), orchestrate on the client instead, or hand it to an application-level gateway (see Model Gateway and Canary Releases).

Common Pitfalls and Troubleshooting ​

PitfallSymptomFix
Missing backendLogs show unable to load 'onnxruntime'Use the full -py3 image, or point --backend-config at the path
Wrong shape definitionsunexpected shape 400dims excludes the batch dim; the exported input names must match the config exactly
Wrong batch axisdynamic batching never kicks inMark the batch axis dynamic at export and keep max_batch_size > 0
Doubled GPU memoryOOM after instance_group count=2Each instance holds its own weights: count × weight memory ≤ usable memory
Queue explosionHigh latency, huge queue_durationConcurrency too high or max_queue_delay too large; reduce it or raise count
Hot reload not workingFile changed, nothing happensCheck --model-control-mode=poll, or switch to explicit and load/unload manually
Bad version directory nameModel shows UNAVAILABLEVersion dirs must be pure numbers (1/, 2/), not v1

Further Reading ​

References ​