Skip to content

Model Serving and Orchestration

At a glance Turning a model from a notebook into a stable production service — the layered architecture of inference serving (API gateway / inference engine / model repository / load balancing / health checks), deployment topologies, mainstream frameworks (Triton / vLLM / TGI / SGLang / Ray Serve), multi-model routing, autoscaling, monitoring metrics, and the MLOps loop.

Model Serving and Orchestration ​

Turning a working model into an online service that can withstand traffic, be observed, and be rolled back takes an entire engineering stack. That is model serving — What Is Inference Acceleration? on this site covers "how to make the model run fast"; this article covers "how to serve the model reliably to the outside world." It sits above Batching and Request Scheduling (how requests are packed into batches), implements the Deployment Design Principles, and wraps the engine case studies like Triton Inference Server and vLLM and PagedAttention.

Where This Article Fits

This article is not a tutorial for any specific framework — it builds the mental model of "what an inference service looks like." After reading it, you should be able to answer: how many layers a service has, what each layer does, which framework to choose, how to monitor, and how to scale. For hands-on framework guides, see the corresponding Triton Inference Server case study; for implementation principles, see Deployment Design Principles.

Concept Definition: The Five Requirements of Model Serving ​

A production-grade inference service must satisfy five things at once — no single one is optional:

RequirementMeaningWhat Happens Without It
AvailabilityThe service stays up and recovers fast when it failsSingle point of failure; one OOM takes down the whole site
ScalabilityAdd GPUs when traffic rises; save money when it fallsCan't handle the daytime load; burns money at night
ObservabilityKnow latency, throughput, error rate, and GPU utilizationBlack-box operations; guesswork when things break
EvolvabilityHot model updates, canary releases, rollbackReleases require downtime; rollback means restarting
Cost controlPer-token cost is measurable and optimizableGPU utilization < 30%, runaway bills

These five are the fundamental difference between serving and "running a bare model script." A bare script validates "does the model work"; serving validates "can the system go live." See Deployment Design Principles.

1. The Layered Architecture of an Inference Service ​

A mature inference service typically has five layers, from the outside in:

text
┌──────────────────────────────────────────────────┐
│  ① API Gateway / Access Layer                    │  ← Auth, rate limiting, routing
├──────────────────────────────────────────────────┤
│  ② Load Balancer / Scheduling Layer              │  ← Request dispatch, load balancing
├──────────────────────────────────────────────────┤
│  ③ Inference Engine / Engine Layer               │  ← vLLM, TensorRT-LLM, TGI
├──────────────────────────────────────────────────┤
│  ④ Model Repository / Repository Layer           │  ← Versioning, hot reload
├──────────────────────────────────────────────────┤
│  ⑤ Hardware / Hardware Layer                     │  ← GPU/CPU/NPU + monitoring
└──────────────────────────────────────────────────┘

① API Gateway (Access Layer) ​

Exposes HTTP/gRPC interfaces to the outside world, handling authentication (API key / OAuth), rate limiting (token bucket / leaky bucket), and routing (by model name / tenant). LLM scenarios additionally require streaming responses (SSE / WebSocket) and long-connection timeouts — generating 2000 tokens can take tens of seconds, and a traditional gateway's default 30s timeout will truncate the response.

② Load Balancer (Scheduling Layer) ​

Dispatches requests to backend inference instances. Common strategies:

StrategyHow It WorksBest For
Round robinDispatch in turnHomogeneous instances and requests
Least connectionsSend to the lowest current concurrencyThe LLM default (request durations vary widely)
Consistent hashingSame prompt goes to the same instancePairs with Batching and Request Scheduling's prefix caching
Route by modelDifferent models to different instance poolsMulti-model serving

The hard part of load balancing for LLM inference: request durations are long and vary wildly (a short prompt takes 0.5s, a long one 10s) — plain round robin piles long requests onto one instance. Least connections + health checks is the de facto standard.

③ Inference Engine (Engine Layer) ​

The layer that actually runs the model — the core of this site. See Inference Engine Comparison for the mainstream engines:

  • vLLM and PagedAttention: the largest open-source community; PagedAttention + continuous batching;
  • TensorRT-LLM: NVIDIA official; FP8 + in-flight batching;
  • TGI: from Hugging Face, closely integrated with the Transformers ecosystem;
  • Triton Inference Server: NVIDIA's multi-backend serving framework for unified management of multiple engines.

Batching and scheduling inside the engine layer are covered in Batching and Request Scheduling; metric definitions in Latency, Throughput, and Concurrency.

④ Model Repository (Repository Layer) ​

Manages model-file versioning, storage, and hot reload. Requirements:

  • Versioning: every release has a unique version number and can be rolled back;
  • Hot reload: load a new model without downtime (Triton's --model-control=explicit + poll);
  • A/B and canary: run two versions at once and split traffic by ratio;
  • Storage formats: LLMs commonly use SafeTensors / GGUF / TensorRT Plan — see vLLM and PagedAttention and llama.cpp and GGUF.

For tooling, MLflow Model Registry is the open-source de facto standard, paired with object storage (S3/OSS) as the model store.

⑤ Hardware (Hardware Layer) ​

GPUs/CPUs/NPUs plus the accompanying drivers, CUDA, and NCCL. Hardware selection: see the Hardware Primer; measured hardware data: Benchmark Data & Tool Profiles. Monitoring at this layer (GPU utilization, memory, temperature) is the basis for scheduling decisions above.

2. Deployment Topologies ​

Inference services come in four typical deployment topologies, with completely different applicability and cost structures.

Single Machine, Single GPU ​

The simplest: one process + one GPU + one model. Suited for prototype validation and low-traffic internal services. The bottleneck is single-GPU memory and compute — a 70B model in FP16 doesn't fit; you must quantize or upgrade the hardware.

Single Machine, Multi-GPU (TP) ​

Multiple GPUs in one machine; Distributed Inference (TP/PP)'s tensor parallelism shards each layer's weights across GPUs. Requires NVLink interconnect (see the Hardware Primer); suited to single-machine inference of 70B+ models. 8×H100 SXM is a common configuration.

Multi-Machine, Multi-GPU (PP / DP) ​

  • Pipeline Parallelism (PP): the model is split by layers across machines — little communication but pipeline bubbles; suited to very large models (405B);
  • Data Parallelism (DP): each machine runs a full model replica with a load balancer in front; suited to scaling throughput (add replicas as traffic grows).

See Distributed Inference (TP/PP).

Serverless / Edge ​

  • Serverless: instances cold-start per request — suited to bursty traffic or low-frequency calls. LLM cold starts are slow (minutes to load a model), so warmup and model caching are required;
  • Edge deployment: phones / car systems / IoT — INT4 quantization + small models are mandatory; see Mobile Deployment and llama.cpp and GGUF.

Bare Serverless Doesn't Suit LLMs

LLM models routinely weigh tens of GB; cold-starting into GPU memory takes tens of seconds to minutes — far beyond the millisecond scale of traditional services. The common compromise: always-on instances + autoscaling (keep a few instances running, scale out as traffic rises) rather than pure per-request cold starts. See Deployment Design Principles.

3. Comparing Mainstream Serving Frameworks ​

Different frameworks serve different purposes — this is not "who replaces whom."

FrameworkPositioningStrengthsBest For
Triton Inference ServerMulti-backend serving frameworkUnified management of TensorRT/PyTorch/ONNX models, dynamic batchingUnified serving for multiple models/engines
vLLM and PagedAttention serverLLM-specialized engine + serverPagedAttention, continuous batching, OpenAI-compatible APIHigh-throughput serving of a single LLM
TGIHugging Face's official LLM serverDeep integration with Transformers/HubUsers of the HF ecosystem
SGLangStructured-generation frameworkRadixAttention prefix caching, JSON/tool callingStructured output, multi-turn dialogue
Ray ServeGeneral ML serving frameworkElastic scaling, multi-model pipelines, Python-nativeComplex orchestration, multi-model pipelines
BentoMLML model packaging and deploymentPackage models into portable imagesCross-environment deployment

Three Rules of Thumb for Choosing a Framework

  1. Single-LLM serving: go straight to vLLM or TensorRT-LLM — HTTP serving and continuous batching built in;
  2. Unified multi-model/multi-engine serving: use Triton, plugging vLLM/TensorRT in as backends;
  3. Complex orchestration (multi-model pipelines, RAG): use Ray Serve or BentoML for the upper-level orchestration, calling vLLM/Triton underneath. Don't reach for Ray Serve on day one — a simple LLM service is one vLLM command away. See Inference Engine Comparison.

4. Multi-Model Routing and Traffic Management ​

Production services often deploy multiple models at once (different sizes, tasks, and versions); the routing layer decides "which model a request goes to."

Routing Strategies ​

StrategyHow It WorksScenario
By model nameRoute by the request's model=llama-3-70bMultiple models coexist (OpenAI-compatible API)
By taskClassification → small model; generation → large modelHeterogeneous tasks
By tenantVIP customers to the large model; free tier to the small oneMulti-tenant SLA
By costSimple requests to the small model; complex ones to the largeCost optimization
A/B canary5% of traffic to the new versionModel releases
CascadeThe small model answers first; escalate to the large model when unsureCost reduction via cascaded inference

Cascade Inference ​

An advanced cost-optimization play: a small model (8B) generates first, and a confidence judgment decides; uncertain requests escalate to a large model (70B) for a re-answer. Overall cost can drop to 20-30% of the original while quality stays close to always-large. See Tuning and Performance Optimization.

5. Autoscaling and Capacity Planning ​

The core of autoscaling is "add or remove inference instances by load." LLM services scale harder than traditional services — each instance loads the model slowly, has a large memory footprint, and needs warmup after startup.

Scaling Metrics ​

MetricMeaningScaling Trigger
QPSRequests per secondScale out above threshold
ConcurrencyRequests currently in flightScale out beyond the per-instance ceiling
GPU utilizationShare of compute usedScale out above 80%; scale in below 30%
P99 latencyTail latencyScale out when the SLA is breached
KV cache usageMemory pressureScale out near the ceiling

Capacity Planning: Deriving from Little's Law ​

By Latency, Throughput, and Concurrency's Little's Law (concurrency = throughput × latency), hardware requirements can be derived from business goals:

text
Goal: 1000 QPS, P99 latency < 2s
→ Max concurrency = 1000 × 2 = 2000
→ One H100 sustains ~200 concurrent (Llama-3-70B INT4)
→ Need 2000 / 200 = 10 H100s
→ Add 30% redundancy = 13 GPUs

Per-GPU concurrency must be measured with Inference Benchmarking in Practice, not guessed. See Latency, Throughput, and Concurrency and Benchmark Data & Tool Profiles.

Three Autoscaling Pitfalls for LLMs

  1. Slow cold starts: from instance pull-up to serving takes tens of seconds to minutes (model loading + warmup) — you can't scale in seconds like a web service; scale ahead of time (based on QPS trend prediction);
  2. Don't scale in too aggressively: in-flight requests must finish — killing instances outright loses requests; set a grace period;
  3. Memory fragmentation: long-running instances fragment their KV cache and throughput decays — restart instances periodically (e.g. every 24h). See Common Pitfalls and Anti-Patterns.

6. Monitoring Metrics and Observability ​

"A service without monitoring is a black box" — monitoring an LLM inference service is more complex than traditional services because business, engine, and hardware layers must be watched at once.

Three-Layer Monitoring ​

LayerMetricsTools
BusinessQPS, P50/P95/P99 latency, error rate, token throughputPrometheus + Grafana
EngineBatch size, KV cache hit rate, queue length, prefix cache hitsvLLM metrics / Triton metrics
HardwareGPU utilization, memory usage, temperature, powerDCGM Exporter

Key Metrics in Detail ​

  • P50/P95/P99 latency: must be split into TTFT and TPOT — TTFT reflects prefill, TPOT reflects decode; the bottlenecks are completely different. Averages alone are masked by long tails. See Latency, Throughput, and Concurrency.
  • GPU utilization: distinguish "SM utilization" from "Tensor Core utilization" — high SM utilization doesn't imply high Tensor Core utilization; the time may be spent on memory access. Pair it with Curated Resources for timeline analysis.
  • KV cache hit rate: the hit rate of Batching and Request Scheduling's prefix caching — directly determines TTFT in multi-tenant scenarios. Below 20% means prompts have little reuse value or the caching policy is wrong.
  • Queue length: the number of requests queued inside the engine — persistently > 0 means saturation; time to scale out.

The Golden Signals of Monitoring

Four signals are mandatory: latency, traffic, errors, saturation — the "four golden signals" of Google SRE apply equally to LLM inference. Any abnormal signal should trigger an alert. See Deployment Design Principles.

7. The MLOps Loop: From Deployment Back to Retraining ​

Model serving is not the end — it's one stage of a loop. The complete MLOps cycle:

text
Train → Evaluate → Deploy → Monitor → Collect production data → Retrain
                      ↑                                            │
                      └────────────────────────────────────────────┘

The inference stage contributes two things to the loop:

  1. The data flywheel: record real user inputs and model outputs (de-identified first), filter high-value samples back into the training set — see Deployment Design Principles;
  2. Drift detection: monitor changes in the input distribution (prompt length, topics) and output distribution (token counts, refusal rate) to trigger retraining. Terminology: Glossary.

Drift is subtler in LLM scenarios — user question styles shift over time (new topics, new instructions), and the model doesn't "throw errors" — it just "answers badly." You need periodic regression on a fixed evaluation set via Inference Benchmarking in Practice.

8. Trade-offs ​

  • Latency vs. cost: small batches give low latency but low GPU utilization and high cost; large batches give high throughput but high latency — interactive scenarios favor latency, batch scenarios favor throughput;
  • Single machine vs. distributed: single machine is simple but capped; distributed scales but needs NVLink and adds scheduling overhead — see Distributed Inference (TP/PP);
  • Self-hosted vs. managed: self-hosting is flexible but operationally heavy; managed (cloud-vendor LLM inference services) is worry-free but pricier and locks you in — see Deployment Design Principles;
  • Precision vs. capacity: FP8/INT4 quantization frees memory for more concurrency, but accuracy loss must be monitored for quality regression.

Further Reading ​

References ​