Appearance
A Brief History of Deployment: From Scripts to the LLM Inference Era
Model deployment has only about ten years of history, yet it has already moved through four eras: hand-rolled scripts → serving frameworks → cloud native and the engine wars → LLM inference. Understanding why each tool exists matters more than memorizing how to use it — today's ONNX Runtime, Triton, vLLM, and KServe are each a direct response to the previous era's pain points. This page breaks the timeline into four eras of tools, pain points, and turning points, taking you from "can use the tools" to "understands the logic behind them."
1. Timeline Overview: Four Eras
text
2013-2016 2016-2018 2018-2022 2022-now
┌─────────────┐ ┌─────────────────┐ ┌───────────────────┐ ┌───────────────────┐
│ Generation 1 │→ │ Generation 2 │→ │ Generation 3 │→ │ Generation 4 │
│ Script era │ │ Serving │ │ Cloud native & │ │ LLM inference era │
│ │ │ framework era │ │ engine wars │ │ │
│ Flask + │ │ TFServing 2016 │ │ ONNX 2017/TensorRT │ │ vLLM 2023 │
│ hand-written │ │ TorchServe 2018 │ │ Triton 2018 │ │ (PagedKV, │
│ batch jobs │ │ Single-machine │ │ K8s/KServe 2020 │ │ continuous batch) │
│ Hand-copied │ │ deployments │ │ Serverless rises │ │ KV Cache opt, │
│ weights │ │ gRPC emerges │ │ │ │ speculative dec, │
│ │ │ │ │ │ │ prefix caching │
└─────────────┘ └─────────────────┘ └───────────────────┘ └───────────────────┘
Pain: not Pain: framework Pain: multi-engine Pain: LLM long
reproducible, lock-in, weak fragmentation, sequences, KV
no concurrency, concurrency low GPU utilization memory blowups,
no monitoring high latencyThe pattern across all four generations: every evolution is a systematic answer to the previous generation's pain.
2. Generation One: The Script Era (2013–2016)
Representative tools: hand-written Flask HTTP endpoints, cron batch jobs, and scp/USB-stick copying of weight files.
The typical workflow: after training, manually copy the .pth or .h5 file to a server, write a few dozen lines of Flask that calls model.load_state_dict() then predict(), and keep it alive with supervisor or nohup. For batch scenarios, teams wrote Shell + Python scripts run nightly by cron.
Pain points of the era:
- Not reproducible: the environment was "whatever happened to be installed on that one machine" — move to another machine and everything breaks.
- No concurrency: Flask threads fighting over one GPU; requests queued and timed out, and batching wasn't even a concept yet.
- No monitoring, no rollback: when it went down, it stayed down, apart from a glance at logs; switching versions meant doing everything over.
- Upgrades were disasters: one framework version bump and the whole dependency tree collapsed; nobody dared touch it.
The turning point: as deep learning grew from a "lab toy" into an online component of search, recommendation, and voice products, the fragility of hand-written scripts became a business bottleneck for the first time. The industry realized: shipping a model requires not "code that runs" but "a system that can be managed."
3. Generation Two: The Serving Framework Era (2016–2018)
Representative tools: TensorFlow Serving (2016), TorchServe (2018), Seldon Core (2018), plus early model-format standardization (ONNX, 2017).
TensorFlow Serving was the first software to productize "model serving": it introduced model version management (the model path is the version), hot-swapping new versions, dynamic batching, and gRPC interfaces, going mainstream alongside TensorFlow in 2018. TorchServe was the PyTorch camp's 2018 answer — official, and it filled the serving gap in the PyTorch ecosystem.
Pain points of the era:
- Framework lock-in: TFServing could only serve TensorFlow models; PyTorch had to wait for TorchServe, and crossing frameworks was extremely costly.
- Single-machine ceilings: early frameworks mostly targeted one machine with several GPUs; traffic distribution across machines and replicas was DIY.
- Deployment ≠ serving: the frameworks only solved the "serving" segment; images, releases, scaling, and monitoring still had to be assembled by hand — far from automated.
The turning point: ONNX (Open Neural Network Exchange) arrived in 2017 — the first decoupling of "model format" from "framework." A single ONNX file could be executed by multiple runtimes, setting the stage for the next era's engine wars. See Model Formats and Conversion for how formats evolved.
4. Generation Three: Cloud Native and the Engine Wars (2018–2022)
Representative tools: ONNX Runtime (2018), NVIDIA TensorRT (long-running iteration), NVIDIA Triton (2018), Kubernetes + KServe (KServe's predecessor KFServing, 2020), Serverless (SageMaker Serverless, Knative).
This generation ran on two parallel tracks:
Track 1: the inference engine arms race. ONNX Runtime quickly captured market share with cross-platform support, graph optimization, and multiple execution backends (CPU/GPU/DirectML and more). TensorRT pushed NVIDIA GPU performance to the limit with layer fusion and FP16/INT8 quantization. Triton took a different path — instead of implementing its own kernels, it unified the scheduling of multi-framework engines (TensorRT, ONNX Runtime, PyTorch, and TensorFlow can all plug in), while providing high-throughput batching and concurrent model management. A comparison of the three approaches: Choosing Frameworks and Platforms.
Track 2: deployment went cloud native. Kubernetes became the deployment substrate, and KServe (formerly KFServing) turned "model serving" into a native K8s resource: one-click multi-replica, autoscaling (HPA), canary releases, and serverless scale-to-zero on demand. Deployment stopped being "ops scripts" and became "declarative resources."
Pain points of the era:
- Multi-engine fragmentation: TensorFlow models, PyTorch models, and legacy models each needed their own serving stack; ops costs exploded.
- Low GPU utilization: fixed replica counts were either saturated at peak or idle in between; serverless and autoscaling emerged in response.
- A gap between training and deployment: getting a model from training to production required manual format conversion and manual environment setup; model registries like MLflow and MLOps pipelines became the standard fix.
The turning point: inference shifted from "getting the model to run" to "saturating the GPU and letting the platform manage itself." Fine-grained performance and cost management became the deployment engineer's core battlefield — a trajectory that continues today. See Performance Optimization and Capacity Planning.
5. Generation Four: The LLM Inference Era (2022–now)
Representative tools: vLLM (2023; paper covered in The vLLM Paper), TGI (Hugging Face), TensorRT-LLM, SGLang, and Triton's LLM backends.
Large language models rewrote the constraints of inference entirely: the models are enormous (7B–70B+ parameters), generations are long (a single request produces thousands of tokens), and VRAM is acutely scarce. Autoregressive decoding means every token requires a forward pass and tokens must be generated serially — the traditional "one request, one batch" static batching simply breaks down.
This generation's core innovations all revolve around "keeping the GPU fully loaded during generation":
| Technique | What it solves | Order of magnitude |
|---|---|---|
| Continuous batching | Static batching waits for the slowest request, leaving the GPU idle | Throughput gains of several times up to an order of magnitude |
| PagedAttention / KV cache management | Fragmentation and waste in long-sequence KV cache memory | Marked VRAM utilization gains — see The vLLM Paper |
| Speculative decoding | Autoregressive token-by-token generation is slow | Roughly 2–3× faster generation |
| Prefix caching | Shared prompt prefixes get recomputed | Large latency drops on cache hits |
Three new variables in LLM inference
For traditional CV/NLP inference, the bottleneck is single forward-pass time. For LLM inference, the bottlenecks are KV cache memory, generation length, and within-batch dynamics. The systemic consequence: the scheduler becomes a core component of the inference system for the first time — "who computes first, how big the batch is, and where the cache goes" directly determines throughput. This is fully developed in LLM Inference Optimization.
Pain points of the era: LLM services were absurdly expensive and slow enough to drive users away; deploying an LLM on traditional engines often left GPU utilization below 20%. The turning point was software-layer integration: all-in-one "engine + scheduler + server" frameworks like vLLM turned LLM deployment from an ops nightmare into a one-line command, and made continuous batching and memory optimization standard industry capabilities.
6. How Deployment Footprints Evolved: From Bare Metal to the Edge
Over the decade, deployment footprints evolved along the axis of "lighter, closer, more elastic":
| Era | Footprint | Representative tech | Strengths | Costs |
|---|---|---|---|---|
| 2015 | Physical/virtual machines | Bare metal + supervisor | Simple | Slow scaling, wasted resources |
| 2017 | Containers | Docker | Reproducible, isolated | Scheduling still manual |
| 2019 | Container orchestration | Kubernetes + KServe | Self-healing, elastic, declarative | High ops complexity |
| 2021 | Serverless | Knative, SageMaker Serverless | Scale on demand, scale to zero, best cost | Cold-start latency |
| 2019– | Edge deployment | TensorRT, OpenVINO, Jetson | Low latency, works offline | Constrained hardware, hard to update |
The trade-offs between edge and serverless are covered in Deployment Architecture Patterns; hands-on edge work in TensorRT and Edge Deployment.
7. Lessons from the Decade
Looking back, four patterns stand out:
- Platformization: from "one script set per model" to "one platform for everything." TFServing → Triton → KServe each turned repeated toil into platform capability.
- Standardization: model formats (ONNX), interface protocols (gRPC/HTTP), image specs (OCI), and orchestration standards (K8s) stacked layer by layer, driving the cost of "swapping engines or machines" toward zero.
- The performance arms race: from graph optimization and quantization to continuous batching, every generation pushed "output per GPU" closer to the limit — and performance engineers have only grown scarcer.
- LLMs redrew the boundaries: schedulers, KV caches, and prefix caching became new infrastructure — when the model itself becomes the bottleneck, the system's fate is decided in the engine and the scheduler.
Why history matters
Among the most common deployment interview questions are "why vLLM instead of Triton?" and "what's the relationship between ONNX and TensorRT?" Answering well doesn't come from memorizing docs — it comes from knowing which era each tool belongs to and which pain point it solved. You now have that background — go to Choosing Frameworks and Platforms and turn it into a concrete selection table.
Further Reading
- What Is Model Deployment? — definitions and the four elements, where history starts
- Choosing Frameworks and Platforms — three generations of engines compared for today's decisions
- LLM Inference Optimization — the core mechanisms of generation four
- PagedAttention: The vLLM Paper — how continuous batching and PagedAttention work
- Deployment Architecture Patterns — a selection framework from bare metal to serverless
- Anatomy of an Inference System — what today's inference systems look like
References
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention (arXiv 2309.06180)
- Orca: A Distributed Serving System for Transformer-Based Generative Models (OSDI 2022, the continuous batching paper)
- Accelerating Large Language Model Decoding with Speculative Sampling (arXiv 2302.01318)
- ONNX project and specification
- NVIDIA TensorRT documentation
- KServe documentation
- TensorFlow Serving documentation