Skip to content

Model Serving and Inference APIs

At a glance Serving is the critical leap that turns a model into a callable API. This article explains what an online inference service really is, how to choose between HTTP and gRPC, how to design inputs and outputs, how to wrap preprocessing and postprocessing, and how to manage the model lifecycle.

Model Serving and Inference APIs ​

One-sentence definition: serving is the engineering process of taking a model that runs fine offline and wrapping it into an API service that business teams can call reliably — any client gets an inference result with a single network request, without caring about weights, frameworks, or hardware.

Industry insight: between "the model runs" and "the model is actually used by the business" sits an entire layer of serving engineering. Plenty of teams get instant results on a laptop, then fall apart in production: timeouts as soon as concurrency rises, request formats that change every few days, model updates that require downtime. The industry consensus is that the inference API is the "product" and the model is merely its core; interface stability, observability, and lifecycle management account for over 70% of the serving effort. This section covers that 70%.

1. What Serving Solves ​

ProblemWithout servingWith serving
DependenciesCallers must install Python, PyTorch, CUDAThey just send HTTP/gRPC requests
ConcurrencySingle process, serial, second-level queuingMultiple instances + load balancing + batching
Model updatesCallers change code and redeployThe serving side rolls out a canary; callers notice nothing
Resource reuseEvery business runs its own copy of the modelOne service shared across businesses
ObservabilityBlack boxMetrics, logs, and tracing all in place

In one line: serving turns the "model" from a research artifact into measurable infrastructure.

2. Protocol Choice: HTTP/REST vs gRPC ​

DimensionHTTP/RESTgRPC
Data formatJSON (human-readable, debug-friendly)Protobuf (binary, compact)
PerformanceHeavy serialization overhead, noticeable on large payloadsSerialization 3-10x faster
StreamingSSE (Server-Sent Events) worksNative bidirectional streaming — the first choice for LLMs
EcosystemCallable from any language or toolRequires generated clients; heavier toolchain
DebuggingWorks with plain curlNeeds tools like grpcurl
Typical use casesGeneral business APIs, gateway integrationHigh-performance internal services, LLM streaming

What people actually do

REST/JSON on the outside (business callers and gateways), gRPC on the inside (between model services and on high-throughput paths). Triton and vLLM both offer both interfaces — don't pick one and stick with it forever.

3. Inference API Design ​

1. Input/output schema: define the contract before writing code ​

A typical contract for an online recommendation-model API:

json
{
  "model": "ctr-v3",
  "inputs": {
    "user_id": "u_10086",
    "item_ids": ["i_1", "i_2", "i_3"],
    "context": { "scene": "home", "hour": 21 }
  },
  "params": { "top_k": 5 }
}
json
{
  "code": 0,
  "data": {
    "scores": [0.91, 0.82, 0.76, 0.61, 0.44]
  },
  "trace_id": "6f8c..."
}

Design principles:

  1. Use stable domain semantics for inputs and outputs (like user_id and top_k), not raw tensors — business callers shouldn't have to understand the model;
  2. Version it: put a version in the URL (/v1/predict) or a model field in the payload — never make callers guess;
  3. Single-item first: online endpoints default to single items; use a dedicated endpoint for batches (see below).

2. Single vs batch ​

  • Single: latency-first, for real-time scenarios;
  • Batch: an array sharing the same schema, computed in one go, throughput-first;
  • Rule: never duplicate logic between the two endpoints — both should route into the same inference core and differ only in the outer loop.

3. Streaming responses (a must for LLMs) ​

Chat-style models generate output token by token (see LLM Inference Optimization). Waiting for the full generation before responding pushes time-to-first-token into the tens of seconds. The standard approach:

text
HTTP + SSE:
  data: {"delta": "Hel"}
  data: {"delta": "lo"}
  data: {"delta": ", "}
  data: {"delta": "world"}
  data: [DONE]

TTFT (time to first token) and TPOT (time per output token) are the core metrics for this kind of API — stop measuring everything as a single "overall latency" figure when you load test.

4. Error code design ​

ScenarioRecommendation
Invalid parameters400 + a clear error message naming the offending fields
Model timeout504 (at the gateway) or 408
Model not ready / overloaded503 + Retry-After
Server-side failure500 + trace_id (makes log lookup easy)

Don't leak internal details in error messages

An error like 500: cuDNN error in MatMul hands attackers your framework and CUDA versions for free. In production, standardize on "code + short message + trace_id"; see Security, Privacy, and Compliance.

4. Preprocessing/Postprocessing: It Must Match Training ​

The easiest way to wreck an inference service — and the hardest failure to debug — is feature definition drift: training normalized with (x - mean)/std but the online preprocessing forgot to subtract the mean; the text classifier was trained with full-width punctuation cleaned but production skipped that step. Quality drops across the board, even though the model itself never changed.

Three iron rules:

  1. Feature engineering code shares one source with training code: extract preprocessing into its own module, have training and inference import the same package, and forbid "rewriting it" for production;
  2. Version it: preprocessing code ships with the model version — when a model rolls back, preprocessing rolls back with it;
  3. Run offline parity checks before launch: feed the same batch of samples through training-side preprocessing and online preprocessing, then compare field by field.

For more pitfalls see Common Pitfalls and Anti-Patterns.

5. Model Lifecycle Management ​

1. Loading strategies ​

text
Eager load  : no delay on the first request, but slow startup (7B weights take tens of seconds to load)
Lazy load   : slow first request, normal afterwards; suits low-traffic services
Warmup      : after loading, run a few dummy requests so kernels and memory layouts are ready (mandatory for TensorRT)

Production advice: loading + warmup must finish before the health check passes. Otherwise every scale-out introduces new instances whose first requests all time out — the number-one cause of scaling incidents.

2. Hot updates and multi-version coexistence ​

  • Multi-version coexistence: one service instance loads both v1 and v2; new requests go to v2 while existing connections finish on v1, for a smooth switchover;
  • Canary releases: route 5% of traffic to the new version first, watch the dashboards, then roll out fully (see Release Strategies: Canary and Rollback);
  • Decouple models from code: keep model files in separate storage (e.g., object storage with hash addressing) so model updates never touch service code.

6. Concurrency, Scheduling, and Request Queues ​

An online service is inherently a queuing system:

text
Request ──► Receive ──► Rate limit (optional) ──► Queue ──► Scheduler (dynamic batching) ──► GPU inference ──► Response
  • Bound your queues: an unbounded queue accumulates millions of requests during a traffic spike and drags out recovery time indefinitely. Set a cap; once it's exceeded, return 503 immediately.
  • Dynamic batching: the scheduler groups requests arriving within a ~10ms window into one batch and computes them together — roughly doubling throughput for a few milliseconds of extra latency. This is the single most important scheduling technique for online services; see Performance Optimization and Capacity Planning.
  • Concurrency model: async non-blocking + multiple instances. Avoid "one thread per request + Python computation under the GIL".

7. A Survey of Serving Frameworks ​

FrameworkOne-line positioningBest for
FastAPILightweight Python API framework; you own preprocessing and orchestrationSmall-to-mid-size teams, fast launches
TorchServeOfficial PyTorch; model packaging and multi-versioning out of the boxPyTorch ecosystems that want official support
Triton Inference ServerMulti-framework, multi-model, high throughput; dynamic batching as a first-class citizenProduction-grade, mixed model portfolios
KServeKubernetes-native, serverless autoscalingCloud-native K8s teams
BentoMLPackages models into a standard artifact (a Bento)Teams that need an engineered packaging pipeline
Ray ServeDistributed with flexible orchestrationComplex services, multi-model compositions

For an in-depth framework comparison see How to Choose a Framework or Platform. For hands-on experience with FastAPI and Triton, see Online Serving with FastAPI + Docker and Multi-Model Serving with NVIDIA Triton.

8. Reliability: Timeouts, Retries, Rate Limiting, Circuit Breaking ​

MechanismPurposeRules of thumb
TimeoutStops callers from waiting foreverSet to 3-5x the model's P99 latency
RetryTolerates transient failuresExponential backoff + jitter; at most 2-3 attempts
Rate limitingProtects the backend from being overwhelmedToken bucket; limit per key and per IP
Circuit breakerFails fast when the backend is brokenBreak for 30s when the consecutive error rate exceeds 50%

Retry storms

Upstream timeouts firing retries + a downstream that happens to be overloaded = avalanche. You must cap retry counts and retry concurrency, and give different callers different quotas, so one failure can't take down the entire cluster.

9. Health Checks and Readiness Probes ​

  • Liveness probe: the process is alive — that's all it checks;
  • Readiness probe: model loaded + warmup complete + able to serve requests; otherwise the instance is pulled from load balancing;
  • Startup probe: gives slow-loading models (like 7B) time to finish loading, so the readiness probe doesn't kill them prematurely.

Example configuration in Kubernetes:

yaml
readinessProbe:
  httpGet: { path: /health/ready, port: 8000 }
  initialDelaySeconds: 60   # wait for the model to finish loading
  periodSeconds: 10
livenessProbe:
  httpGet: { path: /health/live, port: 8000 }
  periodSeconds: 30

Trade-offs ​

DecisionOptionsHow to choose
REST vs gRPCGenerality vs performanceREST externally; gRPC for high-throughput internals
FrameworkLightweight FastAPI vs heavyweight TritonFastAPI for simple single-model cases; Triton for multi-model, high-throughput
Loadingeager vs lazyeager + warmup in production; lazy for low-traffic services
Timeouts and retriesLoose vs strictStrict limits — preventing avalanches comes first
QueuesBounded vs unboundedAlways bounded; 503 as soon as the cap is exceeded

One-line summary: serving "productizes" the model — define the contract first (the API), then guarantee stability (timeouts / rate limits / retries), then manage the lifecycle (loading / warmup / canary). The model itself turns out to be the simplest part.

Further Reading ​

References ​