Appearance
Serving Systems: Clipper / Orca / Nexus
Half of model deployment is "making the model smaller"; the other half is "keeping the service fast and reliable" — and that half belongs to inference serving systems. This walkthrough follows four key papers along the timeline: TensorFlow Serving (the paradigm definer), Clipper (the low-latency serving layer), Orca (iteration-level scheduling, the precursor of continuous batching), and Nexus (GPU sharding and multi-tenancy). By the end you'll see that every feature in today's Triton, KServe, and vLLM traces back to a 2016–2022 paper.
Read Model Serving and Inference APIs and LLM Inference Optimization first to build the conceptual foundation.
Overview Table
| System | Year/Venue | Core contribution | Who inherited its ideas |
|---|---|---|---|
| TensorFlow Serving | 2016 system / 2017 paper (NIPS ML Systems Workshop) | Model versioning, hot loading, dynamic batching, gRPC — the definer of the "serving framework" paradigm | NVIDIA Triton, KServe, TorchServe |
| Clipper | 2017 / NSDI | A general prediction-serving layer: caching, latency-aware batching, adaptive model selection | Model gateways/routers, batching schedulers, A/B experimentation systems |
| Nexus | 2019 / SOSP | GPU cluster inference: executing DNNs as fragments, time-sliced scheduling, multi-tenant sharing | GPU MIG, Triton multi-model concurrency, PD disaggregation (prefill/decode split) |
| Orca | 2022 / OSDI | Iteration-level scheduling + selective batching, making LLM batching actually work | Continuous batching in vLLM, TGI, TensorRT-LLM |
1. TensorFlow Serving: The Paradigm Definer (Olston et al., 2016/2017)
The contribution in one sentence
The first production-grade model serving framework: it hardened the components of "model as a service" — servable versioning, hot loading, dynamic batching, and gRPC interfaces — into a reusable paradigm, defining a layered architecture still in use today.
Background and motivation
Around 2015, every Google model launch relied on hand-rolled scripts: a Flask app here, a full restart on every version bump, a crash under any concurrency. TensorFlow Serving's goal was to turn "shipping a model" into "configuring a servable": the model path is the version — write the config for the new version, hot-load it, traffic switches over automatically, and rollback is available.
The method in brief
- Servables and version management: a servable is a "served model instance," each with a version number (encoded in the model export path); the Manager keeps multiple versions live, and lookups address
(servable, version). - Hot loading and graceful switching: once a new version finishes loading, the switch is atomic; old versions are reclaimed by reference count — no service interruption.
- Dynamic batching: requests queue first and are aggregated into batches by "batch size threshold + maximum wait time" before inference — the knob between throughput and latency.
- Interfaces: gRPC (Protobuf), supporting multiple frameworks (TensorFlow in the paper, extensible by design).
text
Client → gRPC
→ predict request (servable, version, input)
→ Servable Manager (version selection / hot loading)
→ dynamic batcher (accumulate: max_batch_size / batch_timeout)
→ model execution (GPU/CPU)Key results
The paper emphasizes architecture and engineering over benchmarks: the core claim is that the hot path of model lookup and inference was carefully optimized, avoiding the performance traps of naive implementations — and it powered Google's internal multi-tenant hosted service TFS². Its biggest contribution was defining the problem: versioning, batching, and hot loading became the mandatory checklist for every "serving framework" from then on.
Limitations
- Single-framework lock-in (TensorFlow models); multi-framework serving had to wait for Triton.
- Early versions were single-machine designs; cross-machine scaling and elasticity weren't its territory.
What it means today
- Triton directly inherited versioning, dynamic batching, and gRPC, and made "multi-framework backends" its headline feature — see the Triton case study.
- KServe evolved "version = deployment config" into a native K8s resource (canaries and rollbacks done declaratively).
- The high-frequency interview question "why do we need model versioning?" is answered in this very paper.
2. Clipper: A General Low-Latency Prediction Serving Layer (Crankshaw et al., NSDI 2017)
The contribution in one sentence
A general serving layer between applications and ML frameworks: with caching, latency-aware adaptive batching, and model selection, prediction services gain throughput and robustness while meeting latency objectives — a "model serving platform," not "a service for one model."
Background and motivation
TensorFlow Serving answered "how do I serve one model"; Clipper asked a bigger question: when an application calls many models, or when one function has several candidate models, who is the unified entry point? Its goal: "applications write the calling logic once; how the underlying models are deployed, selected, and batched is entirely the serving layer's job."
The method in brief
- Caching: similar/identical queries hit the cache and return directly, skipping inference — a huge win for high-frequency repeated queries like image retrieval and recommendations.
- Latency-aware adaptive batching: no fixed batch size; instead, batch size and wait time adjust dynamically to satisfy a latency objective (e.g. P99 < 100ms).
- Adaptive model selection: maintain multiple models of varying accuracy/speed for the same function and route dynamically based on load and latency budget — high-accuracy models when load is low, fast models when load is high (multi-armed-bandit-style online learning).
- Modular architecture: application layer (client library) → model selection layer → query management (cache/batching) → model execution layer, with pluggable pieces.
Key results
- Across 4 benchmark datasets, throughput and robustness under latency constraints beat naive deployment.
- Against TensorFlow Serving: comparable throughput and latency, while additionally supporting model composition and online learning (caching/model selection brought accuracy and robustness gains).
- The key metric is the latency objective attainment rate: Clipper's scheduler optimizes "meeting P99/tail latency targets," not "average latency" — exactly how SLO-driven capacity planning works today.
Limitations
- Built for fast single-request inference (classification, retrieval), not for LLMs' long autoregressive tasks.
- Model selection requires a set of candidate models with differing accuracy/speed — not every scenario has one.
What it means today
- Model gateways and routing: Clipper's model selection is the intellectual ancestor of today's LLM model gateways, multi-model routing, and A/B traffic splitting.
- Latency objective attainment: P99/SLO — the core metrics of modern capacity planning and observability — were made a scheduler's optimization target by Clipper back in 2017.
- Caching: the LLM era's "prefix caching" is the same idea in a new form (see The vLLM Paper).
3. Nexus: A GPU Cluster Inference Engine (Shen et al., SOSP 2019)
The contribution in one sentence
Reframing GPU-cluster DNN inference as a "fragment + time slice" scheduling problem: instead of running whole models, DNNs execute as fragments scheduled on demand; many video analytics applications share GPUs, pushing throughput near optimal under latency constraints.
Background and motivation
Video analytics (real-time object detection, activity recognition) was AWS's large-scale inference scenario in 2019: dozens to hundreds of cameras, each a continuously processing "application." Giving each application its own GPU explodes cost and leaves utilization in the gutter. Nexus's goal: let one GPU cluster efficiently serve a large number of concurrent video analytics applications while meeting each application's low-latency SLO.
The method in brief
- DNN fragment execution: rather than handing the whole graph to a DNN in one go, video streams are cut into "frame blocks" and the DNN executes at fragment granularity on the GPU — so multiple applications interleave and share a GPU instead of queueing for the whole card.
- Frame Processing Unit: the frame block is the basic scheduling unit, supporting time-sliced sharing — the GPU processes application A's fragments in one slice, then B's in the next.
- Preemptive scheduling: high-priority applications can preempt low-priority ones, protecting latency SLOs.
- Frame sharing/deduplication: when multiple applications analyze the same stream, feature-extraction computations are shared.
Key results
| Metric | Number |
|---|---|
| Throughput (16 GPUs, latency constraint met 99% of the time) | Request processing rate 1.8–12.7× higher than the era's SOTA |
| Utilization (long-running multi-application deployments) | Held near 84% of optimal utilization |
| SLO (100-GPU cluster) | Only 0.27% of requests violated the latency SLO |
Limitations
- Targets streaming video analytics DNNs (CNN-family), not LLM autoregressive generation (which waited for Orca).
- Fragment execution requires model/operator splitting support — significant engineering cost.
What it means today
- GPU time-slicing/memory sharding: Nexus's "carve the GPU into schedulable pieces" is the intellectual precursor of today's MIG (Multi-Instance GPU), Triton multi-model concurrency, and GPU-sharing schedulers on K8s (e.g. time-slicing).
- Fragment execution: the LLM era's PD disaggregation (separate prefill/decode deployment) and chunked prefill are essentially "split one large computation into fragments and schedule the GPU more finely" — Nexus opened this road.
- Multi-tenant SLO guarantees: it provided a quantifiable evaluation framework (SLO violation rate) for "one cluster serving many models."
4. Orca: Iteration-Level Scheduling (Yu et al., OSDI 2022)
The contribution in one sentence
Iteration-level scheduling for Transformer generative models: dropping the batching granularity from "request" to "iteration," combined with selective batching, lifting LLM serving throughput by orders of magnitude — the inventor of continuous batching.
Background and motivation
In 2022, GPT-3-class models were just entering online service, but traditional inference frameworks (built on request-level scheduling) performed terribly on LLMs: one request means dozens to hundreds of iterations (one per generated token), and the framework could only make the whole batch wait for the slowest request:
- Finished requests in a batch can't return early; the GPU waits on the "shortest plank";
- New requests must wait for the entire current batch to end;
- Result: low GPU utilization, poor tail latency, throughput locked by the batching mechanism.
The method in brief
- Iteration-level scheduling: the scheduler re-decides the batch's composition after every iteration (each generated token) — finished requests dequeue immediately, new requests enqueue immediately, and the batch stays full.
- Selective batching: not every operator benefits from batching. Orca batches only the "memory/bandwidth-bound, batching-friendly" operators (matrix multiplications) and keeps autoregression-specific per-token operators unbatched — avoiding the drag of force-batching unfriendly operators.
- Distributed scaling (model parallelism), targeting 10B–100B+ parameter models.
Key results
| Metric | Number |
|---|---|
| Throughput (GPT-3 175B, equal latency) | 36.9× higher than NVIDIA FasterTransformer |
| Latency | Significantly lower than FasterTransformer (no more in-batch mutual drag) |
| Mechanism impact | Iteration-level scheduling became standard equipment in every subsequent LLM inference engine |
Why 36.9× is so extreme
Not because Orca's kernels beat FasterTransformer's, but because of the scheduling mechanism: FasterTransformer holds a fixed batch and waits for the slowest request; Orca keeps the batch full and the GPU always computing. A mechanism-level difference can crush kernel-level optimization by an order of magnitude — the textbook case for "systems papers earn their value in scheduling, not kernels."
Limitations
- KV cache still pre-allocated as contiguous memory — the fragmentation problem remained unsolved, which is precisely the last puzzle piece vLLM's PagedAttention supplied (see The vLLM Paper).
- The scheduling policy is simple (FCFS-style), with no fine-grained priority/preemption design.
What it means today
- Continuous batching is the throughput bedrock of vLLM, TGI, and TensorRT-LLM — and the principle behind parameters like
max_num_seqsin the vLLM case study. - Selection wisdom: scheduling granularity sets the throughput ceiling of a service; kernel optimization is how you scrape toward that ceiling.
- The Orca → vLLM arc is a complete knowledge thread: scheduling solves "stay saturated," memory management solves "fit more requests," and together they stack into 2–4× and beyond.
5. Tying It All Together
text
TFServing (2016) Clipper (2017) Nexus (2019) Orca (2022)
Defined "serving Defined "serving Defined "GPU Defined "LLM
frameworks" layer" sharding" scheduling"
Version/batching/ Cache/batching/ Fragment exec./ Iteration-level
gRPC model selection time-slice/preempt sched./selective batch
│ │ │ │
▼ ▼ ▼ ▼
Triton/KServe Model gateways/ MIG/Triton vLLM/continuous
routing concurrency batching
+ vLLM: merging batching + scheduling + memory management for the LLM eraThe most direct takeaway for today: the three problems you'll hit in deployment — "how do multiple models share one GPU," "how should batching be tuned," and "why is my LLM throughput bad" — each has a first-principles answer in these four papers.
Further Reading
- PagedAttention: The vLLM Paper — the synthesis of Orca scheduling plus memory paging
- Model Serving and Inference APIs — the conceptual foundation of serving frameworks
- Multi-Model Serving with NVIDIA Triton — the TFServing paradigm's modern engineering form
- LLM Serving with vLLM — continuous batching in real configuration
- Parallel and Distributed Inference — what to do when one machine can't fit the model
- Performance Optimization and Capacity Planning — the modern engineering approach to latency objectives
References
- TensorFlow-Serving: Flexible, High-Performance ML Serving (arXiv 1712.06139)
- Clipper: A Low-Latency Online Prediction Serving System (arXiv 1612.03079, NSDI 2017)
- Nexus: A GPU Cluster Engine for Accelerating DNN-based Video Analysis (SOSP 2019)
- Orca: A Distributed Serving System for Transformer-Based Generative Models (OSDI 2022)