Appearance
Model Serving and Orchestration
Turning a working model into an online service that can withstand traffic, be observed, and be rolled back takes an entire engineering stack. That is model serving — What Is Inference Acceleration? on this site covers "how to make the model run fast"; this article covers "how to serve the model reliably to the outside world." It sits above Batching and Request Scheduling (how requests are packed into batches), implements the Deployment Design Principles, and wraps the engine case studies like Triton Inference Server and vLLM and PagedAttention.
Where This Article Fits
This article is not a tutorial for any specific framework — it builds the mental model of "what an inference service looks like." After reading it, you should be able to answer: how many layers a service has, what each layer does, which framework to choose, how to monitor, and how to scale. For hands-on framework guides, see the corresponding Triton Inference Server case study; for implementation principles, see Deployment Design Principles.
Concept Definition: The Five Requirements of Model Serving
A production-grade inference service must satisfy five things at once — no single one is optional:
| Requirement | Meaning | What Happens Without It |
|---|---|---|
| Availability | The service stays up and recovers fast when it fails | Single point of failure; one OOM takes down the whole site |
| Scalability | Add GPUs when traffic rises; save money when it falls | Can't handle the daytime load; burns money at night |
| Observability | Know latency, throughput, error rate, and GPU utilization | Black-box operations; guesswork when things break |
| Evolvability | Hot model updates, canary releases, rollback | Releases require downtime; rollback means restarting |
| Cost control | Per-token cost is measurable and optimizable | GPU utilization < 30%, runaway bills |
These five are the fundamental difference between serving and "running a bare model script." A bare script validates "does the model work"; serving validates "can the system go live." See Deployment Design Principles.
1. The Layered Architecture of an Inference Service
A mature inference service typically has five layers, from the outside in:
text
┌──────────────────────────────────────────────────┐
│ ① API Gateway / Access Layer │ ← Auth, rate limiting, routing
├──────────────────────────────────────────────────┤
│ ② Load Balancer / Scheduling Layer │ ← Request dispatch, load balancing
├──────────────────────────────────────────────────┤
│ ③ Inference Engine / Engine Layer │ ← vLLM, TensorRT-LLM, TGI
├──────────────────────────────────────────────────┤
│ ④ Model Repository / Repository Layer │ ← Versioning, hot reload
├──────────────────────────────────────────────────┤
│ ⑤ Hardware / Hardware Layer │ ← GPU/CPU/NPU + monitoring
└──────────────────────────────────────────────────┘① API Gateway (Access Layer)
Exposes HTTP/gRPC interfaces to the outside world, handling authentication (API key / OAuth), rate limiting (token bucket / leaky bucket), and routing (by model name / tenant). LLM scenarios additionally require streaming responses (SSE / WebSocket) and long-connection timeouts — generating 2000 tokens can take tens of seconds, and a traditional gateway's default 30s timeout will truncate the response.
② Load Balancer (Scheduling Layer)
Dispatches requests to backend inference instances. Common strategies:
| Strategy | How It Works | Best For |
|---|---|---|
| Round robin | Dispatch in turn | Homogeneous instances and requests |
| Least connections | Send to the lowest current concurrency | The LLM default (request durations vary widely) |
| Consistent hashing | Same prompt goes to the same instance | Pairs with Batching and Request Scheduling's prefix caching |
| Route by model | Different models to different instance pools | Multi-model serving |
The hard part of load balancing for LLM inference: request durations are long and vary wildly (a short prompt takes 0.5s, a long one 10s) — plain round robin piles long requests onto one instance. Least connections + health checks is the de facto standard.
③ Inference Engine (Engine Layer)
The layer that actually runs the model — the core of this site. See Inference Engine Comparison for the mainstream engines:
- vLLM and PagedAttention: the largest open-source community; PagedAttention + continuous batching;
- TensorRT-LLM: NVIDIA official; FP8 + in-flight batching;
- TGI: from Hugging Face, closely integrated with the Transformers ecosystem;
- Triton Inference Server: NVIDIA's multi-backend serving framework for unified management of multiple engines.
Batching and scheduling inside the engine layer are covered in Batching and Request Scheduling; metric definitions in Latency, Throughput, and Concurrency.
④ Model Repository (Repository Layer)
Manages model-file versioning, storage, and hot reload. Requirements:
- Versioning: every release has a unique version number and can be rolled back;
- Hot reload: load a new model without downtime (Triton's
--model-control=explicit+ poll); - A/B and canary: run two versions at once and split traffic by ratio;
- Storage formats: LLMs commonly use SafeTensors / GGUF / TensorRT Plan — see vLLM and PagedAttention and llama.cpp and GGUF.
For tooling, MLflow Model Registry is the open-source de facto standard, paired with object storage (S3/OSS) as the model store.
⑤ Hardware (Hardware Layer)
GPUs/CPUs/NPUs plus the accompanying drivers, CUDA, and NCCL. Hardware selection: see the Hardware Primer; measured hardware data: Benchmark Data & Tool Profiles. Monitoring at this layer (GPU utilization, memory, temperature) is the basis for scheduling decisions above.
2. Deployment Topologies
Inference services come in four typical deployment topologies, with completely different applicability and cost structures.
Single Machine, Single GPU
The simplest: one process + one GPU + one model. Suited for prototype validation and low-traffic internal services. The bottleneck is single-GPU memory and compute — a 70B model in FP16 doesn't fit; you must quantize or upgrade the hardware.
Single Machine, Multi-GPU (TP)
Multiple GPUs in one machine; Distributed Inference (TP/PP)'s tensor parallelism shards each layer's weights across GPUs. Requires NVLink interconnect (see the Hardware Primer); suited to single-machine inference of 70B+ models. 8×H100 SXM is a common configuration.
Multi-Machine, Multi-GPU (PP / DP)
- Pipeline Parallelism (PP): the model is split by layers across machines — little communication but pipeline bubbles; suited to very large models (405B);
- Data Parallelism (DP): each machine runs a full model replica with a load balancer in front; suited to scaling throughput (add replicas as traffic grows).
See Distributed Inference (TP/PP).
Serverless / Edge
- Serverless: instances cold-start per request — suited to bursty traffic or low-frequency calls. LLM cold starts are slow (minutes to load a model), so warmup and model caching are required;
- Edge deployment: phones / car systems / IoT — INT4 quantization + small models are mandatory; see Mobile Deployment and llama.cpp and GGUF.
Bare Serverless Doesn't Suit LLMs
LLM models routinely weigh tens of GB; cold-starting into GPU memory takes tens of seconds to minutes — far beyond the millisecond scale of traditional services. The common compromise: always-on instances + autoscaling (keep a few instances running, scale out as traffic rises) rather than pure per-request cold starts. See Deployment Design Principles.
3. Comparing Mainstream Serving Frameworks
Different frameworks serve different purposes — this is not "who replaces whom."
| Framework | Positioning | Strengths | Best For |
|---|---|---|---|
| Triton Inference Server | Multi-backend serving framework | Unified management of TensorRT/PyTorch/ONNX models, dynamic batching | Unified serving for multiple models/engines |
| vLLM and PagedAttention server | LLM-specialized engine + server | PagedAttention, continuous batching, OpenAI-compatible API | High-throughput serving of a single LLM |
| TGI | Hugging Face's official LLM server | Deep integration with Transformers/Hub | Users of the HF ecosystem |
| SGLang | Structured-generation framework | RadixAttention prefix caching, JSON/tool calling | Structured output, multi-turn dialogue |
| Ray Serve | General ML serving framework | Elastic scaling, multi-model pipelines, Python-native | Complex orchestration, multi-model pipelines |
| BentoML | ML model packaging and deployment | Package models into portable images | Cross-environment deployment |
Three Rules of Thumb for Choosing a Framework
- Single-LLM serving: go straight to vLLM or TensorRT-LLM — HTTP serving and continuous batching built in;
- Unified multi-model/multi-engine serving: use Triton, plugging vLLM/TensorRT in as backends;
- Complex orchestration (multi-model pipelines, RAG): use Ray Serve or BentoML for the upper-level orchestration, calling vLLM/Triton underneath. Don't reach for Ray Serve on day one — a simple LLM service is one vLLM command away. See Inference Engine Comparison.
4. Multi-Model Routing and Traffic Management
Production services often deploy multiple models at once (different sizes, tasks, and versions); the routing layer decides "which model a request goes to."
Routing Strategies
| Strategy | How It Works | Scenario |
|---|---|---|
| By model name | Route by the request's model=llama-3-70b | Multiple models coexist (OpenAI-compatible API) |
| By task | Classification → small model; generation → large model | Heterogeneous tasks |
| By tenant | VIP customers to the large model; free tier to the small one | Multi-tenant SLA |
| By cost | Simple requests to the small model; complex ones to the large | Cost optimization |
| A/B canary | 5% of traffic to the new version | Model releases |
| Cascade | The small model answers first; escalate to the large model when unsure | Cost reduction via cascaded inference |
Cascade Inference
An advanced cost-optimization play: a small model (8B) generates first, and a confidence judgment decides; uncertain requests escalate to a large model (70B) for a re-answer. Overall cost can drop to 20-30% of the original while quality stays close to always-large. See Tuning and Performance Optimization.
5. Autoscaling and Capacity Planning
The core of autoscaling is "add or remove inference instances by load." LLM services scale harder than traditional services — each instance loads the model slowly, has a large memory footprint, and needs warmup after startup.
Scaling Metrics
| Metric | Meaning | Scaling Trigger |
|---|---|---|
| QPS | Requests per second | Scale out above threshold |
| Concurrency | Requests currently in flight | Scale out beyond the per-instance ceiling |
| GPU utilization | Share of compute used | Scale out above 80%; scale in below 30% |
| P99 latency | Tail latency | Scale out when the SLA is breached |
| KV cache usage | Memory pressure | Scale out near the ceiling |
Capacity Planning: Deriving from Little's Law
By Latency, Throughput, and Concurrency's Little's Law (concurrency = throughput × latency), hardware requirements can be derived from business goals:
text
Goal: 1000 QPS, P99 latency < 2s
→ Max concurrency = 1000 × 2 = 2000
→ One H100 sustains ~200 concurrent (Llama-3-70B INT4)
→ Need 2000 / 200 = 10 H100s
→ Add 30% redundancy = 13 GPUsPer-GPU concurrency must be measured with Inference Benchmarking in Practice, not guessed. See Latency, Throughput, and Concurrency and Benchmark Data & Tool Profiles.
Three Autoscaling Pitfalls for LLMs
- Slow cold starts: from instance pull-up to serving takes tens of seconds to minutes (model loading + warmup) — you can't scale in seconds like a web service; scale ahead of time (based on QPS trend prediction);
- Don't scale in too aggressively: in-flight requests must finish — killing instances outright loses requests; set a grace period;
- Memory fragmentation: long-running instances fragment their KV cache and throughput decays — restart instances periodically (e.g. every 24h). See Common Pitfalls and Anti-Patterns.
6. Monitoring Metrics and Observability
"A service without monitoring is a black box" — monitoring an LLM inference service is more complex than traditional services because business, engine, and hardware layers must be watched at once.
Three-Layer Monitoring
| Layer | Metrics | Tools |
|---|---|---|
| Business | QPS, P50/P95/P99 latency, error rate, token throughput | Prometheus + Grafana |
| Engine | Batch size, KV cache hit rate, queue length, prefix cache hits | vLLM metrics / Triton metrics |
| Hardware | GPU utilization, memory usage, temperature, power | DCGM Exporter |
Key Metrics in Detail
- P50/P95/P99 latency: must be split into TTFT and TPOT — TTFT reflects prefill, TPOT reflects decode; the bottlenecks are completely different. Averages alone are masked by long tails. See Latency, Throughput, and Concurrency.
- GPU utilization: distinguish "SM utilization" from "Tensor Core utilization" — high SM utilization doesn't imply high Tensor Core utilization; the time may be spent on memory access. Pair it with Curated Resources for timeline analysis.
- KV cache hit rate: the hit rate of Batching and Request Scheduling's prefix caching — directly determines TTFT in multi-tenant scenarios. Below 20% means prompts have little reuse value or the caching policy is wrong.
- Queue length: the number of requests queued inside the engine — persistently > 0 means saturation; time to scale out.
The Golden Signals of Monitoring
Four signals are mandatory: latency, traffic, errors, saturation — the "four golden signals" of Google SRE apply equally to LLM inference. Any abnormal signal should trigger an alert. See Deployment Design Principles.
7. The MLOps Loop: From Deployment Back to Retraining
Model serving is not the end — it's one stage of a loop. The complete MLOps cycle:
text
Train → Evaluate → Deploy → Monitor → Collect production data → Retrain
↑ │
└────────────────────────────────────────────┘The inference stage contributes two things to the loop:
- The data flywheel: record real user inputs and model outputs (de-identified first), filter high-value samples back into the training set — see Deployment Design Principles;
- Drift detection: monitor changes in the input distribution (prompt length, topics) and output distribution (token counts, refusal rate) to trigger retraining. Terminology: Glossary.
Drift is subtler in LLM scenarios — user question styles shift over time (new topics, new instructions), and the model doesn't "throw errors" — it just "answers badly." You need periodic regression on a fixed evaluation set via Inference Benchmarking in Practice.
8. Trade-offs
- Latency vs. cost: small batches give low latency but low GPU utilization and high cost; large batches give high throughput but high latency — interactive scenarios favor latency, batch scenarios favor throughput;
- Single machine vs. distributed: single machine is simple but capped; distributed scales but needs NVLink and adds scheduling overhead — see Distributed Inference (TP/PP);
- Self-hosted vs. managed: self-hosting is flexible but operationally heavy; managed (cloud-vendor LLM inference services) is worry-free but pricier and locks you in — see Deployment Design Principles;
- Precision vs. capacity: FP8/INT4 quantization frees memory for more concurrency, but accuracy loss must be monitored for quality regression.
Further Reading
- Batching and Request Scheduling — how the engine layer packs requests into batches
- Latency, Throughput, and Concurrency — the core metrics a served system must monitor
- The GPU Memory Hierarchy and the Bandwidth Wall — the hard constraint on capacity planning
- Triton Inference Server — the multi-backend serving framework case study
- vLLM and PagedAttention — LLM engine with serving built in
- Distributed Inference (TP/PP) — multi-machine, multi-GPU deployment
- Deployment Design Principles — design principles for serving
- Inference Engine Comparison — framework selection
- Inference Benchmarking in Practice — capacity planning is built on measurement
- Common Pitfalls and Anti-Patterns — common ways serving goes wrong
- Hardware Primer — hardware-layer parameters
- Benchmark Data & Tool Profiles — measured capacity data
- Glossary — definitions of terms used in this article
References
- Triton Inference Server official documentation — the multi-backend inference serving framework
- vLLM official documentation — LLM inference engine and serving
- Ray Serve documentation — general ML serving orchestration framework
- Google SRE Book — The Four Golden Signals — the golden-signals monitoring methodology
- MLflow Model Registry — model versioning and registry