Skip to content

A Brief History of Deployment: From Scripts to the LLM Inference Era

At a glance From 2015's "scripts plus hand-copied weights" to today's vLLM continuous batching and cloud-native inference platforms — how far has model deployment come in a decade? This page walks the timeline through four eras, their key tools, pain points, and turning points, so you understand where today's tools came from.

A Brief History of Deployment: From Scripts to the LLM Inference Era ​

Model deployment has only about ten years of history, yet it has already moved through four eras: hand-rolled scripts → serving frameworks → cloud native and the engine wars → LLM inference. Understanding why each tool exists matters more than memorizing how to use it — today's ONNX Runtime, Triton, vLLM, and KServe are each a direct response to the previous era's pain points. This page breaks the timeline into four eras of tools, pain points, and turning points, taking you from "can use the tools" to "understands the logic behind them."

1. Timeline Overview: Four Eras ​

text
2013-2016          2016-2018             2018-2022                 2022-now
┌─────────────┐  ┌─────────────────┐  ┌───────────────────┐  ┌───────────────────┐
│ Generation 1 │→ │ Generation 2     │→ │ Generation 3       │→ │ Generation 4       │
│ Script era   │  │ Serving          │  │ Cloud native &     │  │ LLM inference era  │
│              │  │ framework era    │  │ engine wars        │  │                    │
│ Flask +      │  │ TFServing 2016   │  │ ONNX 2017/TensorRT │  │ vLLM 2023          │
│ hand-written │  │ TorchServe 2018  │  │ Triton 2018        │  │ (PagedKV,          │
│ batch jobs   │  │ Single-machine   │  │ K8s/KServe 2020    │  │ continuous batch)  │
│ Hand-copied  │  │ deployments      │  │ Serverless rises   │  │ KV Cache opt,      │
│ weights      │  │ gRPC emerges     │  │                    │  │ speculative dec,   │
│              │  │                  │  │                    │  │ prefix caching     │
└─────────────┘  └─────────────────┘  └───────────────────┘  └───────────────────┘
  Pain: not        Pain: framework       Pain: multi-engine    Pain: LLM long
  reproducible,    lock-in, weak         fragmentation,        sequences, KV
  no concurrency,  concurrency           low GPU utilization   memory blowups,
  no monitoring                                                high latency

The pattern across all four generations: every evolution is a systematic answer to the previous generation's pain.

2. Generation One: The Script Era (2013–2016) ​

Representative tools: hand-written Flask HTTP endpoints, cron batch jobs, and scp/USB-stick copying of weight files.

The typical workflow: after training, manually copy the .pth or .h5 file to a server, write a few dozen lines of Flask that calls model.load_state_dict() then predict(), and keep it alive with supervisor or nohup. For batch scenarios, teams wrote Shell + Python scripts run nightly by cron.

Pain points of the era:

  • Not reproducible: the environment was "whatever happened to be installed on that one machine" — move to another machine and everything breaks.
  • No concurrency: Flask threads fighting over one GPU; requests queued and timed out, and batching wasn't even a concept yet.
  • No monitoring, no rollback: when it went down, it stayed down, apart from a glance at logs; switching versions meant doing everything over.
  • Upgrades were disasters: one framework version bump and the whole dependency tree collapsed; nobody dared touch it.

The turning point: as deep learning grew from a "lab toy" into an online component of search, recommendation, and voice products, the fragility of hand-written scripts became a business bottleneck for the first time. The industry realized: shipping a model requires not "code that runs" but "a system that can be managed."

3. Generation Two: The Serving Framework Era (2016–2018) ​

Representative tools: TensorFlow Serving (2016), TorchServe (2018), Seldon Core (2018), plus early model-format standardization (ONNX, 2017).

TensorFlow Serving was the first software to productize "model serving": it introduced model version management (the model path is the version), hot-swapping new versions, dynamic batching, and gRPC interfaces, going mainstream alongside TensorFlow in 2018. TorchServe was the PyTorch camp's 2018 answer — official, and it filled the serving gap in the PyTorch ecosystem.

Pain points of the era:

  • Framework lock-in: TFServing could only serve TensorFlow models; PyTorch had to wait for TorchServe, and crossing frameworks was extremely costly.
  • Single-machine ceilings: early frameworks mostly targeted one machine with several GPUs; traffic distribution across machines and replicas was DIY.
  • Deployment ≠ serving: the frameworks only solved the "serving" segment; images, releases, scaling, and monitoring still had to be assembled by hand — far from automated.

The turning point: ONNX (Open Neural Network Exchange) arrived in 2017 — the first decoupling of "model format" from "framework." A single ONNX file could be executed by multiple runtimes, setting the stage for the next era's engine wars. See Model Formats and Conversion for how formats evolved.

4. Generation Three: Cloud Native and the Engine Wars (2018–2022) ​

Representative tools: ONNX Runtime (2018), NVIDIA TensorRT (long-running iteration), NVIDIA Triton (2018), Kubernetes + KServe (KServe's predecessor KFServing, 2020), Serverless (SageMaker Serverless, Knative).

This generation ran on two parallel tracks:

Track 1: the inference engine arms race. ONNX Runtime quickly captured market share with cross-platform support, graph optimization, and multiple execution backends (CPU/GPU/DirectML and more). TensorRT pushed NVIDIA GPU performance to the limit with layer fusion and FP16/INT8 quantization. Triton took a different path — instead of implementing its own kernels, it unified the scheduling of multi-framework engines (TensorRT, ONNX Runtime, PyTorch, and TensorFlow can all plug in), while providing high-throughput batching and concurrent model management. A comparison of the three approaches: Choosing Frameworks and Platforms.

Track 2: deployment went cloud native. Kubernetes became the deployment substrate, and KServe (formerly KFServing) turned "model serving" into a native K8s resource: one-click multi-replica, autoscaling (HPA), canary releases, and serverless scale-to-zero on demand. Deployment stopped being "ops scripts" and became "declarative resources."

Pain points of the era:

  • Multi-engine fragmentation: TensorFlow models, PyTorch models, and legacy models each needed their own serving stack; ops costs exploded.
  • Low GPU utilization: fixed replica counts were either saturated at peak or idle in between; serverless and autoscaling emerged in response.
  • A gap between training and deployment: getting a model from training to production required manual format conversion and manual environment setup; model registries like MLflow and MLOps pipelines became the standard fix.

The turning point: inference shifted from "getting the model to run" to "saturating the GPU and letting the platform manage itself." Fine-grained performance and cost management became the deployment engineer's core battlefield — a trajectory that continues today. See Performance Optimization and Capacity Planning.

5. Generation Four: The LLM Inference Era (2022–now) ​

Representative tools: vLLM (2023; paper covered in The vLLM Paper), TGI (Hugging Face), TensorRT-LLM, SGLang, and Triton's LLM backends.

Large language models rewrote the constraints of inference entirely: the models are enormous (7B–70B+ parameters), generations are long (a single request produces thousands of tokens), and VRAM is acutely scarce. Autoregressive decoding means every token requires a forward pass and tokens must be generated serially — the traditional "one request, one batch" static batching simply breaks down.

This generation's core innovations all revolve around "keeping the GPU fully loaded during generation":

TechniqueWhat it solvesOrder of magnitude
Continuous batchingStatic batching waits for the slowest request, leaving the GPU idleThroughput gains of several times up to an order of magnitude
PagedAttention / KV cache managementFragmentation and waste in long-sequence KV cache memoryMarked VRAM utilization gains — see The vLLM Paper
Speculative decodingAutoregressive token-by-token generation is slowRoughly 2–3× faster generation
Prefix cachingShared prompt prefixes get recomputedLarge latency drops on cache hits

Three new variables in LLM inference

For traditional CV/NLP inference, the bottleneck is single forward-pass time. For LLM inference, the bottlenecks are KV cache memory, generation length, and within-batch dynamics. The systemic consequence: the scheduler becomes a core component of the inference system for the first time — "who computes first, how big the batch is, and where the cache goes" directly determines throughput. This is fully developed in LLM Inference Optimization.

Pain points of the era: LLM services were absurdly expensive and slow enough to drive users away; deploying an LLM on traditional engines often left GPU utilization below 20%. The turning point was software-layer integration: all-in-one "engine + scheduler + server" frameworks like vLLM turned LLM deployment from an ops nightmare into a one-line command, and made continuous batching and memory optimization standard industry capabilities.

6. How Deployment Footprints Evolved: From Bare Metal to the Edge ​

Over the decade, deployment footprints evolved along the axis of "lighter, closer, more elastic":

EraFootprintRepresentative techStrengthsCosts
2015Physical/virtual machinesBare metal + supervisorSimpleSlow scaling, wasted resources
2017ContainersDockerReproducible, isolatedScheduling still manual
2019Container orchestrationKubernetes + KServeSelf-healing, elastic, declarativeHigh ops complexity
2021ServerlessKnative, SageMaker ServerlessScale on demand, scale to zero, best costCold-start latency
2019–Edge deploymentTensorRT, OpenVINO, JetsonLow latency, works offlineConstrained hardware, hard to update

The trade-offs between edge and serverless are covered in Deployment Architecture Patterns; hands-on edge work in TensorRT and Edge Deployment.

7. Lessons from the Decade ​

Looking back, four patterns stand out:

  1. Platformization: from "one script set per model" to "one platform for everything." TFServing → Triton → KServe each turned repeated toil into platform capability.
  2. Standardization: model formats (ONNX), interface protocols (gRPC/HTTP), image specs (OCI), and orchestration standards (K8s) stacked layer by layer, driving the cost of "swapping engines or machines" toward zero.
  3. The performance arms race: from graph optimization and quantization to continuous batching, every generation pushed "output per GPU" closer to the limit — and performance engineers have only grown scarcer.
  4. LLMs redrew the boundaries: schedulers, KV caches, and prefix caching became new infrastructure — when the model itself becomes the bottleneck, the system's fate is decided in the engine and the scheduler.

Why history matters

Among the most common deployment interview questions are "why vLLM instead of Triton?" and "what's the relationship between ONNX and TensorRT?" Answering well doesn't come from memorizing docs — it comes from knowing which era each tool belongs to and which pain point it solved. You now have that background — go to Choosing Frameworks and Platforms and turn it into a concrete selection table.

Further Reading ​

References ​