Appearance
Model Serving and Inference APIs
One-sentence definition: serving is the engineering process of taking a model that runs fine offline and wrapping it into an API service that business teams can call reliably — any client gets an inference result with a single network request, without caring about weights, frameworks, or hardware.
Industry insight: between "the model runs" and "the model is actually used by the business" sits an entire layer of serving engineering. Plenty of teams get instant results on a laptop, then fall apart in production: timeouts as soon as concurrency rises, request formats that change every few days, model updates that require downtime. The industry consensus is that the inference API is the "product" and the model is merely its core; interface stability, observability, and lifecycle management account for over 70% of the serving effort. This section covers that 70%.
1. What Serving Solves
| Problem | Without serving | With serving |
|---|---|---|
| Dependencies | Callers must install Python, PyTorch, CUDA | They just send HTTP/gRPC requests |
| Concurrency | Single process, serial, second-level queuing | Multiple instances + load balancing + batching |
| Model updates | Callers change code and redeploy | The serving side rolls out a canary; callers notice nothing |
| Resource reuse | Every business runs its own copy of the model | One service shared across businesses |
| Observability | Black box | Metrics, logs, and tracing all in place |
In one line: serving turns the "model" from a research artifact into measurable infrastructure.
2. Protocol Choice: HTTP/REST vs gRPC
| Dimension | HTTP/REST | gRPC |
|---|---|---|
| Data format | JSON (human-readable, debug-friendly) | Protobuf (binary, compact) |
| Performance | Heavy serialization overhead, noticeable on large payloads | Serialization 3-10x faster |
| Streaming | SSE (Server-Sent Events) works | Native bidirectional streaming — the first choice for LLMs |
| Ecosystem | Callable from any language or tool | Requires generated clients; heavier toolchain |
| Debugging | Works with plain curl | Needs tools like grpcurl |
| Typical use cases | General business APIs, gateway integration | High-performance internal services, LLM streaming |
What people actually do
REST/JSON on the outside (business callers and gateways), gRPC on the inside (between model services and on high-throughput paths). Triton and vLLM both offer both interfaces — don't pick one and stick with it forever.
3. Inference API Design
1. Input/output schema: define the contract before writing code
A typical contract for an online recommendation-model API:
json
{
"model": "ctr-v3",
"inputs": {
"user_id": "u_10086",
"item_ids": ["i_1", "i_2", "i_3"],
"context": { "scene": "home", "hour": 21 }
},
"params": { "top_k": 5 }
}json
{
"code": 0,
"data": {
"scores": [0.91, 0.82, 0.76, 0.61, 0.44]
},
"trace_id": "6f8c..."
}Design principles:
- Use stable domain semantics for inputs and outputs (like
user_idandtop_k), not raw tensors — business callers shouldn't have to understand the model; - Version it: put a version in the URL (
/v1/predict) or amodelfield in the payload — never make callers guess; - Single-item first: online endpoints default to single items; use a dedicated endpoint for batches (see below).
2. Single vs batch
- Single: latency-first, for real-time scenarios;
- Batch: an array sharing the same schema, computed in one go, throughput-first;
- Rule: never duplicate logic between the two endpoints — both should route into the same inference core and differ only in the outer loop.
3. Streaming responses (a must for LLMs)
Chat-style models generate output token by token (see LLM Inference Optimization). Waiting for the full generation before responding pushes time-to-first-token into the tens of seconds. The standard approach:
text
HTTP + SSE:
data: {"delta": "Hel"}
data: {"delta": "lo"}
data: {"delta": ", "}
data: {"delta": "world"}
data: [DONE]TTFT (time to first token) and TPOT (time per output token) are the core metrics for this kind of API — stop measuring everything as a single "overall latency" figure when you load test.
4. Error code design
| Scenario | Recommendation |
|---|---|
| Invalid parameters | 400 + a clear error message naming the offending fields |
| Model timeout | 504 (at the gateway) or 408 |
| Model not ready / overloaded | 503 + Retry-After |
| Server-side failure | 500 + trace_id (makes log lookup easy) |
Don't leak internal details in error messages
An error like 500: cuDNN error in MatMul hands attackers your framework and CUDA versions for free. In production, standardize on "code + short message + trace_id"; see Security, Privacy, and Compliance.
4. Preprocessing/Postprocessing: It Must Match Training
The easiest way to wreck an inference service — and the hardest failure to debug — is feature definition drift: training normalized with (x - mean)/std but the online preprocessing forgot to subtract the mean; the text classifier was trained with full-width punctuation cleaned but production skipped that step. Quality drops across the board, even though the model itself never changed.
Three iron rules:
- Feature engineering code shares one source with training code: extract preprocessing into its own module, have training and inference import the same package, and forbid "rewriting it" for production;
- Version it: preprocessing code ships with the model version — when a model rolls back, preprocessing rolls back with it;
- Run offline parity checks before launch: feed the same batch of samples through training-side preprocessing and online preprocessing, then compare field by field.
For more pitfalls see Common Pitfalls and Anti-Patterns.
5. Model Lifecycle Management
1. Loading strategies
text
Eager load : no delay on the first request, but slow startup (7B weights take tens of seconds to load)
Lazy load : slow first request, normal afterwards; suits low-traffic services
Warmup : after loading, run a few dummy requests so kernels and memory layouts are ready (mandatory for TensorRT)Production advice: loading + warmup must finish before the health check passes. Otherwise every scale-out introduces new instances whose first requests all time out — the number-one cause of scaling incidents.
2. Hot updates and multi-version coexistence
- Multi-version coexistence: one service instance loads both v1 and v2; new requests go to v2 while existing connections finish on v1, for a smooth switchover;
- Canary releases: route 5% of traffic to the new version first, watch the dashboards, then roll out fully (see Release Strategies: Canary and Rollback);
- Decouple models from code: keep model files in separate storage (e.g., object storage with hash addressing) so model updates never touch service code.
6. Concurrency, Scheduling, and Request Queues
An online service is inherently a queuing system:
text
Request ──► Receive ──► Rate limit (optional) ──► Queue ──► Scheduler (dynamic batching) ──► GPU inference ──► Response- Bound your queues: an unbounded queue accumulates millions of requests during a traffic spike and drags out recovery time indefinitely. Set a cap; once it's exceeded, return 503 immediately.
- Dynamic batching: the scheduler groups requests arriving within a ~10ms window into one batch and computes them together — roughly doubling throughput for a few milliseconds of extra latency. This is the single most important scheduling technique for online services; see Performance Optimization and Capacity Planning.
- Concurrency model: async non-blocking + multiple instances. Avoid "one thread per request + Python computation under the GIL".
7. A Survey of Serving Frameworks
| Framework | One-line positioning | Best for |
|---|---|---|
| FastAPI | Lightweight Python API framework; you own preprocessing and orchestration | Small-to-mid-size teams, fast launches |
| TorchServe | Official PyTorch; model packaging and multi-versioning out of the box | PyTorch ecosystems that want official support |
| Triton Inference Server | Multi-framework, multi-model, high throughput; dynamic batching as a first-class citizen | Production-grade, mixed model portfolios |
| KServe | Kubernetes-native, serverless autoscaling | Cloud-native K8s teams |
| BentoML | Packages models into a standard artifact (a Bento) | Teams that need an engineered packaging pipeline |
| Ray Serve | Distributed with flexible orchestration | Complex services, multi-model compositions |
For an in-depth framework comparison see How to Choose a Framework or Platform. For hands-on experience with FastAPI and Triton, see Online Serving with FastAPI + Docker and Multi-Model Serving with NVIDIA Triton.
8. Reliability: Timeouts, Retries, Rate Limiting, Circuit Breaking
| Mechanism | Purpose | Rules of thumb |
|---|---|---|
| Timeout | Stops callers from waiting forever | Set to 3-5x the model's P99 latency |
| Retry | Tolerates transient failures | Exponential backoff + jitter; at most 2-3 attempts |
| Rate limiting | Protects the backend from being overwhelmed | Token bucket; limit per key and per IP |
| Circuit breaker | Fails fast when the backend is broken | Break for 30s when the consecutive error rate exceeds 50% |
Retry storms
Upstream timeouts firing retries + a downstream that happens to be overloaded = avalanche. You must cap retry counts and retry concurrency, and give different callers different quotas, so one failure can't take down the entire cluster.
9. Health Checks and Readiness Probes
- Liveness probe: the process is alive — that's all it checks;
- Readiness probe: model loaded + warmup complete + able to serve requests; otherwise the instance is pulled from load balancing;
- Startup probe: gives slow-loading models (like 7B) time to finish loading, so the readiness probe doesn't kill them prematurely.
Example configuration in Kubernetes:
yaml
readinessProbe:
httpGet: { path: /health/ready, port: 8000 }
initialDelaySeconds: 60 # wait for the model to finish loading
periodSeconds: 10
livenessProbe:
httpGet: { path: /health/live, port: 8000 }
periodSeconds: 30Trade-offs
| Decision | Options | How to choose |
|---|---|---|
| REST vs gRPC | Generality vs performance | REST externally; gRPC for high-throughput internals |
| Framework | Lightweight FastAPI vs heavyweight Triton | FastAPI for simple single-model cases; Triton for multi-model, high-throughput |
| Loading | eager vs lazy | eager + warmup in production; lazy for low-traffic services |
| Timeouts and retries | Loose vs strict | Strict limits — preventing avalanches comes first |
| Queues | Bounded vs unbounded | Always bounded; 503 as soon as the cap is exceeded |
One-line summary: serving "productizes" the model — define the contract first (the API), then guarantee stability (timeouts / rate limits / retries), then manage the lifecycle (loading / warmup / canary). The model itself turns out to be the simplest part.
Further Reading
- Online Serving with FastAPI + Docker — a minimal but complete service, from zero to one
- Multi-Model Serving with NVIDIA Triton — production-grade multi-model serving in practice
- Deployment Architecture Patterns — how online serving differs from batch and serverless
- Performance Optimization and Capacity Planning — dynamic batching and concurrency models
- Common Pitfalls and Anti-Patterns — classic traps such as feature definition drift