Appearance
A Brief History
Inference acceleration did not explode overnight. Over the past fifteen years it went through four paradigm shifts — from CPU serial to GPU parallel, from static graph optimization to dynamic batching, from dense compute to memory-access-aware, and from single-machine serving to distributed LLM engines — and none of them was merely "switching tools." Each was the result of hardware capability, model architecture, and workload characteristics resonating together. Understanding this line of evolution matters far more than memorizing a string of engine version numbers — it shows you where today's LLM inference engineering sits in history, and where the next wave might come from.
1. A Three-Minute Overview
text
Before 2010s CPU inference era: serial compute, rule-based encoding -> inference = "running the weights"
2012-2017 GPU inference era: cuDNN + operator libraries + graph fusion -> rise of TensorRT/OpenVINO
2017-2022 Transformer era: BERT/GPT change the workload -> Triton/serving matures
2022-present LLM inference revolution: autoregressive + KV cache + batching -> vLLM/SGLang/speculative decodingThe essence of every shift was "the bottleneck of inference optimization moved from one dimension to another": compute -> memory access -> scheduling -> arithmetic intensity. Every part of an LLM inference engineer's daily work today can be located on this timeline as "the bottleneck of the previous generation it solved."
2. The CPU Inference Era (Before the 2010s)
Inference = Running a Serial Forward on a CPU
Before the 2010s, "model deployment" was almost equivalent to "running a forward pass on a CPU." Models of that era were mainly SVMs, decision trees, shallow MLPs, and small CNNs (like LeNet); their compute demands were modest, and a serial CPU forward sufficed. OpenVINO — born as Intel's DLDT and officially renamed in 2018 — is a product of this generation. Its core positioning was "optimizing model inference on Intel CPUs/iGPUs," squeezing CPU inference dry through operator fusion, INT8 quantization, and graph optimization.
Inference engineering of this generation focused on two things:
- Operator optimization: hand-written SIMD (AVX2/AVX512) to make a single operator faster on a CPU.
- Graph optimization: constant folding, operator fusion, dead-code elimination — rearranging the whole computation graph into a "CPU-friendly" form.
OpenVINO, the early ONNX Runtime, and TensorFlow Serving (2016) all belong to this generation. See OpenVINO and CPU Inference and ONNX Runtime: Cross-Platform.
The Limits of CPU Inference
The ceiling of CPU inference is set by two facts:
- A low compute ceiling: a 32-core Xeon peaks at about 1 TFLOP (FP32), yet training AlexNet in 2012 already required 500+ GFLOPS — the CPU simply could not feed a CNN.
- Low parallelism: CPUs are designed for low-latency serial tasks and are ill-suited to the matrix parallelism of GEMM.
When CNNs ignited deep learning in 2012, CPU inference immediately hit the compute wall, and inference engineering was pushed onto the GPU.
3. The GPU Inference Era (2012–2017)
2012: AlexNet Ignites GPU Training and Indirectly Drives GPU Inference
In 2012 AlexNet cut ImageNet top-5 error from 26.2% to 15.3%. The key was not the model itself (CNNs had existed since LeNet in 1998) but training on two GTX 580 GPUs — the first proof that "GPUs suit neural networks." GPU adoption on the training side quickly carried over to inference: since training used GPUs, using GPUs for inference followed naturally.
2014–2015: cuDNN Makes GPU Inference Engineering-Driven
In 2014 NVIDIA released cuDNN v2, turning convolutions, pooling, activations, and other operators into a highly optimized GPU library — from then on, inference engineers no longer needed to write CUDA; calling cuDNN APIs delivered 80%+ of GPU performance. VGG (2014), GoogLeNet (2014), and ResNet (2015) appeared in succession, models grew deeper and wider, and CPU inference exited the mainstream entirely.
cuDNN's key contribution was the im2col + GEMM convolution implementation — turning convolution into matrix multiplication to exploit the GPU's Tensor Cores. This looks obvious today but was revolutionary at the time, and it laid the foundation for TensorRT.
2015–2017: TensorRT Turns Inference Engineering into a Product
NVIDIA released TensorRT (v1) in 2015, adding INT8 quantization in v2 in 2016. TensorRT's core positioning: compile the computation graph exported from training frameworks (TensorFlow/PyTorch/Caffe) into a highly optimized inference engine for the GPU. Its three pillars:
- Graph-level fusion: Conv+BN+ReLU, MatMul+Add+GELU automatically fused into a single kernel.
- INT8 calibration: calibrate quantization scales with a small number of samples to quantize FP32 models to INT8, a 2-4x speedup.
- Kernel auto-tuning: pick the optimal kernel for the same operator under different batches/shapes.
TensorRT turned "inference acceleration" from "hand-writing CUDA" into "calling an engine" — and inference engineer became a distinct job title. See TensorRT and GPU Inference.
The Same Period: ONNX and ONNX Runtime (2017–2019)
In 2017 Microsoft and Facebook jointly released ONNX (Open Neural Network Exchange), aiming to unify the model format — the training side (PyTorch/TensorFlow/PaddlePaddle) could all export ONNX, and the inference side could load it with any engine. In 2019 Microsoft released ONNX Runtime to run ONNX models on CPUs, GPUs, and various accelerators. The difference between ONNX Runtime and TensorRT: the former is cross-hardware, the latter is NVIDIA-only; the former is more general, the latter more extreme. See ONNX Runtime: Cross-Platform.
4. The Transformer Era (2017–2022)
2017: The Transformer Appears and Changes the Inference Workload
In 2017 Vaswani et al. published Attention Is All You Need, and the Transformer was born. Its profound impact on inference engineering would take five more years to fully show, but several traits already planted the seeds:
- Attention is memory-access-dense: the intermediate matrices of Q×K^T×V are far larger than a convolution's feature maps, and memory access became the new bottleneck.
- Autoregressive generation: decoder models generate one token at a time, completely different from a CNN's "one forward, one result" — this later became the core pain point of LLM inference.
2018–2019: BERT/GPT Open the Era of Large NLP Models
In 2018 BERT (340M parameters) swept 11 NLP benchmarks, and in 2019 GPT-2 (1.5B parameters) was released. These two models were the first time inference engineering faced models "in the hundreds of millions to billions of parameters" — single-GPU memory began to tighten, and quantization, distillation, and tensor parallelism entered the inference engineering vocabulary.
But the "large models" of this period were still relatively small — BERT-large has only 340M parameters, 680 MB even in FP16, comfortably fitting in a 12 GB GPU. Real memory pressure did not arrive until GPT-3 (175B) in 2020, and GPT-3 was closed-source at the time, so the community had not yet realized the special nature of LLM inference.
2019–2020: Triton Inference Server and the Productization of Serving
In 2018–2019 NVIDIA introduced the Triton Inference Server (originally TensorRT Inference Server), turning "inference serving" into a product. Triton's core value:
- Multi-model management: one serving process hosts multiple models, scheduled across GPU resources.
- Multi-framework support: TensorRT, PyTorch, ONNX Runtime, and a Python backend coexist.
- Dynamic batching: the serving side automatically batches arriving requests, trading latency for short requests against throughput for long ones.
Triton upgraded "model serving" from "wrapping a model in FastAPI" to "industrial-grade inference serving," and it remains the core of NVIDIA's inference stack. See Triton Inference Server.
The focus of inference engineering in this period was: keep the model running stably as an online service — multi-replica, canary release, A/B testing, drift monitoring. Model serving took shape as an independent field; see Model Serving and Orchestration.
5. The LLM Inference Revolution (2022–Present)
2022: ChatGPT Ignites LLM Inference Demand
OpenAI released ChatGPT in November 2022, reaching 100 million users in two months. This was not just a product event but a watershed for inference engineering — suddenly every company had to "deploy a conversational LLM," yet HuggingFace Transformers ran Llama-2-70B at only ~20 tokens/s for a single request, entirely unable to support real traffic.
The fundamental differences between LLM inference and traditional model inference (see Inference vs. Training vs. Fine-Tuning):
- Autoregressive: generating one token requires a full forward pass, unlike a CNN that produces a result in one shot.
- Enormous KV cache: a 70B model with 4k of context needs ~8 GB of KV cache per request, and memory fragmentation is severe.
- Variable request lengths: user prompts range from dozens to thousands of tokens, and static batching wastes heavily.
- Low arithmetic intensity: at batch=1 decode, arithmetic intensity is ~1 FLOP/byte, far below the GPU roofline, with GPU utilization under 5%.
These four traits combined to make "deploying an LLM with transformers" the industry's most painful problem in 2022 — and vLLM was born to solve it.
2023: The vLLM and PagedAttention Revolution
In June 2023 UC Berkeley released vLLM and the PagedAttention paper Efficient Memory Management for LLM Serving with PagedAttention. This was the tipping point of LLM inference engineering — the core idea borrows virtual memory from operating systems:
- PagedAttention: split the KV cache into fixed-size "pages" (blocks), allocate on demand, eliminate memory fragmentation, and lift GPU memory utilization from ~20% to ~90%.
- Continuous batching: requests join and leave the batch dynamically, long and short requests coexist, and throughput improves 10-20x.
Measured on Llama-2-70B, vLLM delivers roughly 14-24x the throughput of HuggingFace Transformers — the key leap that took LLM inference from "runs" to "ships." See vLLM and PagedAttention.
After vLLM, the open-source community followed quickly: HuggingFace shipped TGI (Text Generation Inference), LMSYS shipped SGLang, and TensorRT-LLM added continuous batching. Continuous batching + PagedAttention became the de facto standard for LLM inference engines.
2023: TensorRT-LLM Ships and INT4 Quantization Matures
In the second half of 2023 NVIDIA officially released TensorRT-LLM — extending TensorRT's graph-fusion capability to LLMs, with tensor parallelism, pipeline parallelism, INT4/INT8/FP8 quantization, and in-flight batching (similar to continuous batching). TensorRT-LLM typically beats vLLM's throughput by 30-50% on H100 (NVIDIA's optimizations are more aggressive) but is harder to deploy and has a smaller community. See TensorRT-LLM.
In the same period, GPTQ (ICLR 2023) and AWQ (MLSys 2024) made INT4 weight quantization practical for LLMs:
- GPTQ: layer-by-layer weight quantization based on second-order information, compressing a 70B model from 140 GB to 35 GB so it fits on a single 80 GB A100.
- AWQ: activation-aware weight quantization, preserving the weight channels that are sensitive to activation outliers, losing far less accuracy than naive INT4.
These two methods made "INT4 quantization for deployment" standard equipment for LLM inference. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
2024: SGLang, FlashInfer, EAGLE, and FP8 on H100
2024 was the year LLM inference engineering went from "usable" to "extreme":
- SGLang: from UC Berkeley's LMSYS, introducing RadixAttention (prefix-sharing KV cache) and structured-generation acceleration (compressed FSM), 2-5x faster than vLLM in multi-turn dialogue and function-calling scenarios.
- FlashInfer: an operator library from the same team, unifying attention's prefill/decode into one efficient kernel, widely adopted by vLLM/SGLang/TensorRT-LLM.
- EAGLE (ICML 2024): speculative decoding engineered — a small draft head guesses tokens and the large model verifies them in parallel, 2-3x speedup without changing the output distribution.
- FP8 goes mainstream: the H100's Transformer Engine made FP8 inference practical — less accuracy loss than INT8 and double the compute/memory traffic of FP16, becoming the preferred precision for LLM inference on H100.
This year, operator-level optimization (FlashInfer) and system-level optimization (SGLang's RadixAttention, EAGLE's speculative decoding) became the main lines, and the marginal returns of plain graph fusion began to decline. See Kernel Fusion and Custom Kernels and Speculative Decoding and Medusa/EAGLE.
2025: EAGLE-3, DeepSeek MTP, and FP4 on Blackwell
The main threads of 2025:
- EAGLE-3 / SpecForge: speculative decoding iterated to a third generation, pairing tree-shaped drafting (guessing multiple branches at once) with dynamic depth — measured 2-6x speedup on Llama-3-70B and gradually becoming the default configuration for large-model inference.
- DeepSeek-V3 MTP (Multi-Token Prediction): DeepSeek-V3 introduced an auxiliary head at training time that "predicts multiple tokens at once," so inference natively supports speculative decoding with no external draft model needed — opening the new paradigm of "optimizing for inference at training time."
- FP4 / Blackwell: NVIDIA Blackwell (B200) supports FP4 precision, doubling compute and memory gains over FP8, so a 70B-parameter model can be deployed on a single B200 in FP4 — but FP4 quantization loses more accuracy than FP8, and calibration techniques are still evolving.
- Heterogeneous inference: CPU (Intel AMX, Apple Neural Engine) + GPU + NPU (Huawei Ascend) collaborative inference became a focus for on-device and domestic-deployment scenarios; see Mobile Deployment and Hardware Primer.
6. Two Threads Running Through It All
Looking back over fifteen years, every shift can be attributed to two threads:
Thread one: the bottleneck moved from compute to memory access, then to scheduling. In the 2010s the CPU inference bottleneck was compute (serial); in 2012–2017 the GPU inference bottleneck was graph fusion and operators; in 2017–2022 the Transformer inference bottleneck was attention's memory access; after 2022 the LLM inference bottleneck became batching scheduling and the KV cache. Every generation of engineers solved the bottleneck left behind by the previous generation, then hit a new one.
Thread two: optimization layers permeate from the bottom up. Every shift started at the model layer (quantization/distillation), then spread to the operator layer (FlashAttention/FlashInfer), the graph layer (TensorRT fusion), and the system layer (vLLM batching). The real leap is never one layer but the stacking of all four — today's vLLM + AWQ + FlashAttention-3 + EAGLE-3 combination is the layered product of fifteen years of engineering wisdom.
Three reminders from history
- Engines expire, principles do not: a 2018 TensorRT tutorial reads today with every concrete API outdated, but the principles of "operator fusion + INT8 calibration + graph rearrangement" still hold. Learn principles, not just APIs.
- Benchmark numbers expire: any "vLLM is N times faster than X" figure is a snapshot taken on a specific piece of hardware, a specific model, and a specific workload — check the date before quoting it. Pages in this site's case-studies chapter carry a
dataAsOftag. - The next leap may come from the algorithm side: DeepSeek-V3's MTP shows that "optimizing for inference at training time" may be the next wave — inference engineers can no longer know only engines; they must also understand how model architectures affect inference-friendliness.
Further Reading
- What Is Inference Acceleration? — the conceptual foundation for this page's conclusions
- Anatomy of the Overall Architecture — the complete skeleton of an inference system today
- Inference vs. Training vs. Fine-Tuning — why LLM inference and training are two different fields
- TensorRT and GPU Inference — the benchmark of the GPU inference era
- vLLM and PagedAttention — the tipping point of the LLM inference revolution
- Speculative Decoding and Medusa/EAGLE — the newest thread of 2024-2025
- Classic Papers in Depth — the primary literature at the key nodes of this history
References
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) — the start of the Transformer
- Kwon et al. Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023) — the vLLM paper, the landmark of the LLM inference revolution
- Frantar et al. GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers (ICLR 2023) — the engineering of INT4 LLM quantization
- Lin et al. AWQ: Activation-Aware Weight Quantization (MLSys 2024) — activation-aware weight quantization
- Dao et al. FlashAttention (NeurIPS 2022) — the benchmark of operator-level optimization
- Li et al. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (ICML 2024) — the engineering of speculative decoding
- DeepSeek-AI. DeepSeek-V3 Technical Report (2024) — MTP and training-side inference optimization
- NVIDIA. TensorRT Developer Guide — official documentation of TensorRT's evolution
- NVIDIA. TensorRT-LLM — the open-source repo of the LLM inference engine
- ONNX Runtime Documentation — the evolution of ONNX and cross-hardware inference