Skip to content

Serving Systems: Clipper / Orca / Nexus

At a glance Four key papers on inference serving systems: TensorFlow Serving defined the model serving paradigm, Clipper introduced model selection and adaptive batching, Orca's iteration-level scheduling preceded continuous batching, and Nexus pushed GPU sharding to its limits.

Serving Systems: Clipper / Orca / Nexus ​

Half of model deployment is "making the model smaller"; the other half is "keeping the service fast and reliable" — and that half belongs to inference serving systems. This walkthrough follows four key papers along the timeline: TensorFlow Serving (the paradigm definer), Clipper (the low-latency serving layer), Orca (iteration-level scheduling, the precursor of continuous batching), and Nexus (GPU sharding and multi-tenancy). By the end you'll see that every feature in today's Triton, KServe, and vLLM traces back to a 2016–2022 paper.

Read Model Serving and Inference APIs and LLM Inference Optimization first to build the conceptual foundation.

Overview Table ​

SystemYear/VenueCore contributionWho inherited its ideas
TensorFlow Serving2016 system / 2017 paper (NIPS ML Systems Workshop)Model versioning, hot loading, dynamic batching, gRPC — the definer of the "serving framework" paradigmNVIDIA Triton, KServe, TorchServe
Clipper2017 / NSDIA general prediction-serving layer: caching, latency-aware batching, adaptive model selectionModel gateways/routers, batching schedulers, A/B experimentation systems
Nexus2019 / SOSPGPU cluster inference: executing DNNs as fragments, time-sliced scheduling, multi-tenant sharingGPU MIG, Triton multi-model concurrency, PD disaggregation (prefill/decode split)
Orca2022 / OSDIIteration-level scheduling + selective batching, making LLM batching actually workContinuous batching in vLLM, TGI, TensorRT-LLM

1. TensorFlow Serving: The Paradigm Definer (Olston et al., 2016/2017) ​

The contribution in one sentence ​

The first production-grade model serving framework: it hardened the components of "model as a service" — servable versioning, hot loading, dynamic batching, and gRPC interfaces — into a reusable paradigm, defining a layered architecture still in use today.

Background and motivation ​

Around 2015, every Google model launch relied on hand-rolled scripts: a Flask app here, a full restart on every version bump, a crash under any concurrency. TensorFlow Serving's goal was to turn "shipping a model" into "configuring a servable": the model path is the version — write the config for the new version, hot-load it, traffic switches over automatically, and rollback is available.

The method in brief ​

  • Servables and version management: a servable is a "served model instance," each with a version number (encoded in the model export path); the Manager keeps multiple versions live, and lookups address (servable, version).
  • Hot loading and graceful switching: once a new version finishes loading, the switch is atomic; old versions are reclaimed by reference count — no service interruption.
  • Dynamic batching: requests queue first and are aggregated into batches by "batch size threshold + maximum wait time" before inference — the knob between throughput and latency.
  • Interfaces: gRPC (Protobuf), supporting multiple frameworks (TensorFlow in the paper, extensible by design).
text
Client → gRPC
        → predict request (servable, version, input)
        → Servable Manager (version selection / hot loading)
        → dynamic batcher (accumulate: max_batch_size / batch_timeout)
        → model execution (GPU/CPU)

Key results ​

The paper emphasizes architecture and engineering over benchmarks: the core claim is that the hot path of model lookup and inference was carefully optimized, avoiding the performance traps of naive implementations — and it powered Google's internal multi-tenant hosted service TFS². Its biggest contribution was defining the problem: versioning, batching, and hot loading became the mandatory checklist for every "serving framework" from then on.

Limitations ​

  • Single-framework lock-in (TensorFlow models); multi-framework serving had to wait for Triton.
  • Early versions were single-machine designs; cross-machine scaling and elasticity weren't its territory.

What it means today ​

  • Triton directly inherited versioning, dynamic batching, and gRPC, and made "multi-framework backends" its headline feature — see the Triton case study.
  • KServe evolved "version = deployment config" into a native K8s resource (canaries and rollbacks done declaratively).
  • The high-frequency interview question "why do we need model versioning?" is answered in this very paper.

2. Clipper: A General Low-Latency Prediction Serving Layer (Crankshaw et al., NSDI 2017) ​

The contribution in one sentence ​

A general serving layer between applications and ML frameworks: with caching, latency-aware adaptive batching, and model selection, prediction services gain throughput and robustness while meeting latency objectives — a "model serving platform," not "a service for one model."

Background and motivation ​

TensorFlow Serving answered "how do I serve one model"; Clipper asked a bigger question: when an application calls many models, or when one function has several candidate models, who is the unified entry point? Its goal: "applications write the calling logic once; how the underlying models are deployed, selected, and batched is entirely the serving layer's job."

The method in brief ​

  • Caching: similar/identical queries hit the cache and return directly, skipping inference — a huge win for high-frequency repeated queries like image retrieval and recommendations.
  • Latency-aware adaptive batching: no fixed batch size; instead, batch size and wait time adjust dynamically to satisfy a latency objective (e.g. P99 < 100ms).
  • Adaptive model selection: maintain multiple models of varying accuracy/speed for the same function and route dynamically based on load and latency budget — high-accuracy models when load is low, fast models when load is high (multi-armed-bandit-style online learning).
  • Modular architecture: application layer (client library) → model selection layer → query management (cache/batching) → model execution layer, with pluggable pieces.

Key results ​

  • Across 4 benchmark datasets, throughput and robustness under latency constraints beat naive deployment.
  • Against TensorFlow Serving: comparable throughput and latency, while additionally supporting model composition and online learning (caching/model selection brought accuracy and robustness gains).
  • The key metric is the latency objective attainment rate: Clipper's scheduler optimizes "meeting P99/tail latency targets," not "average latency" — exactly how SLO-driven capacity planning works today.

Limitations ​

  • Built for fast single-request inference (classification, retrieval), not for LLMs' long autoregressive tasks.
  • Model selection requires a set of candidate models with differing accuracy/speed — not every scenario has one.

What it means today ​

  • Model gateways and routing: Clipper's model selection is the intellectual ancestor of today's LLM model gateways, multi-model routing, and A/B traffic splitting.
  • Latency objective attainment: P99/SLO — the core metrics of modern capacity planning and observability — were made a scheduler's optimization target by Clipper back in 2017.
  • Caching: the LLM era's "prefix caching" is the same idea in a new form (see The vLLM Paper).

3. Nexus: A GPU Cluster Inference Engine (Shen et al., SOSP 2019) ​

The contribution in one sentence ​

Reframing GPU-cluster DNN inference as a "fragment + time slice" scheduling problem: instead of running whole models, DNNs execute as fragments scheduled on demand; many video analytics applications share GPUs, pushing throughput near optimal under latency constraints.

Background and motivation ​

Video analytics (real-time object detection, activity recognition) was AWS's large-scale inference scenario in 2019: dozens to hundreds of cameras, each a continuously processing "application." Giving each application its own GPU explodes cost and leaves utilization in the gutter. Nexus's goal: let one GPU cluster efficiently serve a large number of concurrent video analytics applications while meeting each application's low-latency SLO.

The method in brief ​

  • DNN fragment execution: rather than handing the whole graph to a DNN in one go, video streams are cut into "frame blocks" and the DNN executes at fragment granularity on the GPU — so multiple applications interleave and share a GPU instead of queueing for the whole card.
  • Frame Processing Unit: the frame block is the basic scheduling unit, supporting time-sliced sharing — the GPU processes application A's fragments in one slice, then B's in the next.
  • Preemptive scheduling: high-priority applications can preempt low-priority ones, protecting latency SLOs.
  • Frame sharing/deduplication: when multiple applications analyze the same stream, feature-extraction computations are shared.

Key results ​

MetricNumber
Throughput (16 GPUs, latency constraint met 99% of the time)Request processing rate 1.8–12.7× higher than the era's SOTA
Utilization (long-running multi-application deployments)Held near 84% of optimal utilization
SLO (100-GPU cluster)Only 0.27% of requests violated the latency SLO

Limitations ​

  • Targets streaming video analytics DNNs (CNN-family), not LLM autoregressive generation (which waited for Orca).
  • Fragment execution requires model/operator splitting support — significant engineering cost.

What it means today ​

  • GPU time-slicing/memory sharding: Nexus's "carve the GPU into schedulable pieces" is the intellectual precursor of today's MIG (Multi-Instance GPU), Triton multi-model concurrency, and GPU-sharing schedulers on K8s (e.g. time-slicing).
  • Fragment execution: the LLM era's PD disaggregation (separate prefill/decode deployment) and chunked prefill are essentially "split one large computation into fragments and schedule the GPU more finely" — Nexus opened this road.
  • Multi-tenant SLO guarantees: it provided a quantifiable evaluation framework (SLO violation rate) for "one cluster serving many models."

4. Orca: Iteration-Level Scheduling (Yu et al., OSDI 2022) ​

The contribution in one sentence ​

Iteration-level scheduling for Transformer generative models: dropping the batching granularity from "request" to "iteration," combined with selective batching, lifting LLM serving throughput by orders of magnitude — the inventor of continuous batching.

Background and motivation ​

In 2022, GPT-3-class models were just entering online service, but traditional inference frameworks (built on request-level scheduling) performed terribly on LLMs: one request means dozens to hundreds of iterations (one per generated token), and the framework could only make the whole batch wait for the slowest request:

  • Finished requests in a batch can't return early; the GPU waits on the "shortest plank";
  • New requests must wait for the entire current batch to end;
  • Result: low GPU utilization, poor tail latency, throughput locked by the batching mechanism.

The method in brief ​

  • Iteration-level scheduling: the scheduler re-decides the batch's composition after every iteration (each generated token) — finished requests dequeue immediately, new requests enqueue immediately, and the batch stays full.
  • Selective batching: not every operator benefits from batching. Orca batches only the "memory/bandwidth-bound, batching-friendly" operators (matrix multiplications) and keeps autoregression-specific per-token operators unbatched — avoiding the drag of force-batching unfriendly operators.
  • Distributed scaling (model parallelism), targeting 10B–100B+ parameter models.

Key results ​

MetricNumber
Throughput (GPT-3 175B, equal latency)36.9× higher than NVIDIA FasterTransformer
LatencySignificantly lower than FasterTransformer (no more in-batch mutual drag)
Mechanism impactIteration-level scheduling became standard equipment in every subsequent LLM inference engine

Why 36.9× is so extreme

Not because Orca's kernels beat FasterTransformer's, but because of the scheduling mechanism: FasterTransformer holds a fixed batch and waits for the slowest request; Orca keeps the batch full and the GPU always computing. A mechanism-level difference can crush kernel-level optimization by an order of magnitude — the textbook case for "systems papers earn their value in scheduling, not kernels."

Limitations ​

  • KV cache still pre-allocated as contiguous memory — the fragmentation problem remained unsolved, which is precisely the last puzzle piece vLLM's PagedAttention supplied (see The vLLM Paper).
  • The scheduling policy is simple (FCFS-style), with no fine-grained priority/preemption design.

What it means today ​

  • Continuous batching is the throughput bedrock of vLLM, TGI, and TensorRT-LLM — and the principle behind parameters like max_num_seqs in the vLLM case study.
  • Selection wisdom: scheduling granularity sets the throughput ceiling of a service; kernel optimization is how you scrape toward that ceiling.
  • The Orca → vLLM arc is a complete knowledge thread: scheduling solves "stay saturated," memory management solves "fit more requests," and together they stack into 2–4× and beyond.

5. Tying It All Together ​

text
TFServing (2016)     Clipper (2017)       Nexus (2019)        Orca (2022)
Defined "serving     Defined "serving     Defined "GPU        Defined "LLM
frameworks"          layer"               sharding"           scheduling"
Version/batching/    Cache/batching/      Fragment exec./     Iteration-level
gRPC                 model selection      time-slice/preempt  sched./selective batch
    │                  │                    │                  │
    ▼                  ▼                    ▼                  ▼
 Triton/KServe      Model gateways/      MIG/Triton          vLLM/continuous
                    routing              concurrency         batching
 + vLLM: merging batching + scheduling + memory management for the LLM era

The most direct takeaway for today: the three problems you'll hit in deployment — "how do multiple models share one GPU," "how should batching be tuned," and "why is my LLM throughput bad" — each has a first-principles answer in these four papers.

Further Reading ​

References ​