Appearance
What Is Inference Acceleration?
One-sentence definition: inference acceleration is all the engineering that turns trained model weights into a production system that can serve real traffic within latency and cost constraints — training answers "can the model do it," inference answers "can the model be used, afforded, and kept stable."
The most "inference-like" code you have ever written is probably a single line, model.generate(inputs). Unpacked, that line is not computing gradients — it is running one forward pass with a fixed set of parameters to produce an output. The entire discipline of inference engineering starts from the tension behind that line: training counts how many FLOPs you need; inference counts how expensive memory access is, how tight batching is, and how long users wait. One line of model.generate runs Llama-2-70B on HuggingFace Transformers at roughly 20 tokens/s for a single request; the same weights reach 2000+ tokens/s of aggregate throughput on vLLM — a 100x gap that lives not in the model but in the system.
1. Definition: Three Levels of Understanding
1.1 The Practical Level: What It Does
An inference system takes input (prompts / images / features) and produces output (tokens / labels / scores). What makes it distinct is that it is an "online service":
Training: data + labels --> model weights (offline, batch, maximize total throughput)
Inference: model weights + input --> output (online, concurrent, per-request latency x total throughput)Training is "teaching the model in the lab"; inference is "putting the graduated model on the job." Every scenario where "the model is trained and must serve real traffic" — chat APIs, search ranking, recommendation recall, image recognition, speech transcription — is inference engineering's home turf. An inference engineer's daily work is not tuning loss functions but wrestling with latency, GPU memory, throughput, batching, and quantization.
1.2 The Academic Level: A Computation-Graph Definition
From the computation-graph perspective, inference is running a single forward pass over the computation graph — no gradients recorded, no parameters updated:
Given trained parameters θ* and an input x, inference computes y = f(x; θ*), where f is the model definition and θ* is the fixed set of weight tensors. The compute can be characterized exactly as FLOPs = 2 × P × S (P = parameter count, S = sequence length, for a Transformer), and the result is reproducible for the same (θ*, x).
The power of this definition is that it yields three quantifiable objectives for inference optimization:
| Objective | Meaning | Measured by |
|---|---|---|
| Latency | Time for one request from input to first / last token | TTFT, TPOT, E2E latency (see Latency, Throughput, and Concurrency) |
| Throughput | Requests or tokens processed per unit time | tokens/s, QPS, requests/min |
| Memory | GPU memory occupied by weights + activations + KV cache | GB (see The GPU Memory Hierarchy and the Bandwidth Wall) |
Together they form the inference impossible triangle — low latency, high throughput, and small memory usually cannot all be had at once; every optimization trades among the three. See The Roofline Model and Compute Analysis.
1.3 The Mathematical Level: The Memory-Bound View
Mathematically, the core question of inference optimization is not "compute faster" but "feed the compute fast enough":
The arithmetic intensity of one Transformer decode step is roughly
2 × batch × seq_len × hidden_dim(FLOPs) /2 × batch × seq_len × hidden_dim × bytes_per_param(bytes of memory access). At batch=1 and bytes_per_param=2 (FP16), arithmetic intensity is about 1 FLOP/byte — far below the GPU's roofline knee (≈100 FLOPs/byte on an A100), so single-request LLM inference is memory-bound.
This is the single most important fact in inference engineering: single-request LLM decode is not computing — it is waiting for the GPU memory to move weights and the KV cache to the SMs. Every LLM inference optimization — batching, quantization, KV cache compression, kernel fusion — is fundamentally about raising arithmetic intensity. See The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis.
Why is inference "hard"?
The root of training's difficulty is sheer compute and hard convergence; the root of inference's difficulty is extremely uneven compute distribution — the same model has arithmetic intensity ~1 when decoding a single request (memory-bound), ~100 during prefill (compute-bound), and is compute-bound again under batching. One system must feed all three workloads well at once, and that is exactly why inference engineering is "messier" than training engineering.
2. Inference vs. Training: Compute Distribution, Memory Access, and Batching
Inference and training both "run a neural network forward," but engineering-wise they are almost two different fields. The three most important differences:
2.1 Compute Distribution: Training Is Dense and Uniform, Inference Sparse and Uneven
One epoch of training is "one forward plus one backward per sample," backward costs about 2x forward, and every sample is treated equally — compute is dense and uniform, and GPU utilization can be stably pushed to 50-60% MFU.
Inference is the opposite. LLM inference has two phases (see Inference vs. Training vs. Fine-Tuning):
- Prefill phase (processing the prompt): compute-dense, high arithmetic intensity even at batch=1, GPU utilization can reach 60%+.
- Decode phase (generating tokens): each step computes only one token, arithmetic intensity ~1 at batch=1, and GPU utilization is often under 5%.
So for the same Llama-2-70B, prefilling a 1k-token prompt might take only 200 ms, while decoding 200 tokens takes 4 seconds — decode is the latency bottleneck, and decode is almost entirely waiting on memory.
2.2 Memory Access: Training Reuses Weights, Inference Re-Reads Them Every Time
During training, weights are read once per batch, and a large batch size raises the compute/memory-access ratio. In the inference decode phase, every token requires reading the entire set of weights once — a 70B model in FP16 is 140 GB, and generating a single token means moving 140 GB from HBM to the SMs. On an A100's 80 GB with 2 TB/s of bandwidth, the memory-access floor for one token is 70 ms. This is why the token/s ceiling of single-request LLM inference is roughly HBM bandwidth / model size — almost independent of compute.
2.3 Batching: Training Is Static, Inference Is Continuous
Training batch size is fixed within a step (static batching) and can run into the hundreds or thousands. Inference requests arrive at any time and vary in length — with static batching you either wait to fill a batch (latency explodes) or force truncation (experience collapses). This is the core problem that vLLM's PagedAttention and continuous batching solves: let requests join and leave the batch dynamically, packing "N requests of different lengths" onto the same GPU to compute together, raising decode arithmetic intensity from ~1 to ~50+.
The full training-vs-inference comparison is in Inference vs. Training vs. Fine-Tuning.
3. The Four Layers of Optimization (Model / Operator / Graph / System)
Inference optimization is not a single trick but a four-level tree. Any "inference acceleration" technique belongs to one of these four layers:
text
┌──────────┬───────────────────────────────────────────────────────────────────────┐
│ System │ batching · scheduling · KV cache · routing · multi-model │
│ │ e.g. vLLM continuous batching, PagedAttention │
├──────────┼───────────────────────────────────────────────────────────────────────┤
│ Graph │ operator fusion · constant folding · dead-code elim · memory planning │
│ │ e.g. TensorRT graph fusion, ONNX graph optimization │
├──────────┼───────────────────────────────────────────────────────────────────────┤
│ Operator │ FlashAttention · handwritten CUDA kernels · operator fusion │
│ │ e.g. FlashAttention-2/3, FlashInfer │
├──────────┼───────────────────────────────────────────────────────────────────────┤
│ Model │ quantization · pruning · distillation · sparsification │
│ │ e.g. GPTQ/AWQ INT4, SmoothQuant, distillation │
└──────────┴───────────────────────────────────────────────────────────────────────┘
▲ upper layers: bigger gains but a higher bar; lower layers: direct gains but need hardware knowledge ▼
optimize bottom-up, but in practice all four layers are stackedModel Layer: Make the Model Itself Smaller
- Quantization: FP16 → INT8/INT4/FP8, directly cutting weights and memory traffic by 2-4x — the highest-ROI optimization. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
- Pruning: remove unimportant weights/channels/heads; structured pruning genuinely saves compute, unstructured pruning needs sparse-kernel support. See Pruning and Sparsification.
- Distillation: use a large model to teach a small one, letting the small model approach the large model's quality. See Knowledge Distillation.
Operator Layer: Make a Single Operator Fast
- FlashAttention: compute attention's intermediate matrices in SRAM tiles instead of shuttling them back and forth through HBM — a 2-4x speedup for a single operator, now standard equipment for LLM inference. See Kernel Fusion and Custom Kernels.
- Handwritten CUDA/Triton kernels: custom kernels shaped to your own model; vLLM, SGLang, and FlashInfer all use them heavily.
- Operator fusion: merge several adjacent operators into one kernel, reducing kernel launches and memory traffic.
Graph Layer: Reschedule the Whole Computation Graph
- Operator fusion (graph level): identify fusable patterns at the graph level (Conv+BN+ReLU, MatMul+Add+GELU), done automatically by TensorRT and ONNX Runtime.
- Constant folding: compute constants at compile time (e.g. fold
2 × 3 × xinto6 × x). - Memory planning: reorder operators and reuse activation memory to push peak memory down. See Computation Graph Optimization.
System Layer: Feed Many Requests Together
- Continuous batching: let requests of different lengths join and leave the batch dynamically — the core mechanism of vLLM/TGI/SGLang.
- KV cache management: PagedAttention pages the KV cache like virtual memory, lifting GPU memory utilization from 20% to 90%+.
- Speculative decoding: a small/draft model guesses a few tokens, the large model verifies them in parallel — 2-6x speedup. See Speculative Decoding and Medusa/EAGLE.
- Multi-model routing: simple requests go to a small model, complex ones to a large model, cutting overall cost. See Model Serving and Orchestration.
Stacking all four layers is what production looks like
A real production system does not "pick one layer" — it stacks all four. A typical Llama-3-70B service: H100 FP8 at the bottom (hardware layer) → AWQ INT4 quantization (model layer) → FlashAttention-3 + FlashInfer (operator layer) → TensorRT-LLM graph fusion (graph layer) → continuous batching + PagedAttention + EAGLE-3 speculative decoding (system layer). Stacked, a single H100 node goes from HuggingFace's ~20 tokens/s to ~3000 tokens/s — a 150x gap that is all engineering.
4. A Minimal Inference System: Handwritten forward + batched forward + KV cache
Strip away all the wrapping from vLLM and transformers, and a runnable LLM inference loop takes only 60 lines of Python. In this section we hand-write it to see the three levels of inference engineering: single-request forward, batched forward, and autoregressive generation with a KV cache.
4.1 Single Forward: The Plainest Inference
python
import torch
import torch.nn.functional as F
# Assume `model` is a Transformer with weights already loaded (any framework works)
model = load_model("llama-7b", dtype=torch.float16, device="cuda")
input_ids = torch.tensor([[1, 1500, 2500, 980]], device="cuda") # one prompt
# The plainest inference: one forward, take logits, argmax the next token
with torch.no_grad():
logits = model(input_ids) # [1, seq_len, vocab_size]
next_token = logits[0, -1].argmax(-1) # take the logits at the last position
print(next_token) # that is one greedy decode stepNothing is optimized yet, but it already exposes the core problem: every forward recomputes attention over the whole sequence — at token 2 it recomputes the first token, at token 3 it recomputes the first two — O(N²) redundant compute. Generating 200 tokens on a 1000-token prompt means computing attention 1000 + 1001 + ... + 1200 ≈ 220k times, while the genuinely new computation needed is only 200 × (1000 + 200) ≈ 240k — the redundancy is close to half.
4.2 Batching: Compute Several Requests Together
python
# Three requests of different lengths, padded to equal length
prompts = [
[1, 1500, 2500, 980], # length 4
[1, 980, 2333, 880, 1024, 55], # length 6
[1, 733], # length 2
]
max_len = max(len(p) for p in prompts)
# left padding (the attention mask hides pad tokens from the model)
input_ids = torch.full((3, max_len), pad_id, device="cuda")
attention_mask = torch.zeros(3, max_len, device="cuda")
for i, p in enumerate(prompts):
input_ids[i, max_len - len(p):] = p
attention_mask[i, max_len - len(p):] = 1
with torch.no_grad():
logits = model(input_ids, attention_mask=attention_mask)
next_tokens = logits[:, -1, :].argmax(-1) # each request gets its own next tokenAt batch=3, arithmetic intensity triples — but the price is padding waste and having to wait for the slowest request. This is the pain point of static batching, and the problem continuous batching solves.
4.3 KV Cache: Avoid Recomputation
python
# Key idea: let the model cache each layer's K and V; the next step computes only the new token
past_key_values = None
generated = list(prompts[0]) # copy the first prompt as the generated sequence
for _ in range(20): # generate 20 tokens
input_ids = torch.tensor([[generated[-1]]], device="cuda") # feed only the new token
with torch.no_grad():
out = model(input_ids, past_key_values=past_key_values, use_cache=True)
logits = out.logits[:, -1, :]
past_key_values = out.past_key_values # accumulate the KV cache
next_token = logits.argmax(-1).item()
generated.append(next_token)With a KV cache, each step computes the forward of a single token — compute drops by O(seq_len), but memory access becomes reading the entire KV cache each step. This is the most critical data structure in LLM inference, and PagedAttention is precisely a paged management scheme built on top of it.
Actually run it once
All three snippets run in any GPU environment. Try them with transformers' LlamaForCausalLM: add use_cache=False/True, batch=1/4/16 in turn, and record the tokens/s and peak memory at each step — you will find the KV cache is the lifeline of LLM inference and batch size is the lifeline of throughput. The full hands-on walkthrough is in Deploy an Inference Service from Scratch.
5. Why Inference Optimization "Works": Three Empirical Milestones
"Inference optimization speeds things up" is not a marketing slogan — it is a fact you can verify against public data. Three landmark pieces of evidence:
5.1 TensorRT vs. Native PyTorch: 2-4x
TensorRT is NVIDIA's inference engine, applying graph fusion + operator fusion + precision tuning to the same model. On ResNet-50, TensorRT INT8 typically delivers 2-4x the throughput of PyTorch FP32; on BERT-large, TensorRT delivers about 3x the throughput of PyTorch FP16. This speedup requires no model changes and no new hardware — pure engineering. See TensorRT and GPU Inference.
5.2 vLLM PagedAttention: 10-20x
In June 2023, UC Berkeley's vLLM paper Efficient Memory Management for LLM Serving with PagedAttention appeared. Its core idea is managing the KV cache in pages, the way an operating system manages virtual memory, lifting GPU memory utilization from ~20% to ~90% while supporting continuous batching. Measured on Llama-2-70B, vLLM delivers roughly 14-24x the throughput of HuggingFace Transformers (FP16) — the tipping point for LLM inference engineering. See vLLM and PagedAttention.
5.3 EAGLE-3 Speculative Decoding: 2-6x
In 2024-2025, speculative decoding moved from theory to engineering. The EAGLE series uses a small "draft head" to guess a few tokens, then the large model verifies them in parallel and accepts the verified tokens in one shot. EAGLE-3 measured 2-6x speedup on Llama-3-70B without changing the output distribution — an optimization that truly has it both ways. See Speculative Decoding and Medusa/EAGLE.
Stack the three milestones: native PyTorch (1x) → TensorRT (3x) → vLLM (30x) → + EAGLE-3 (90x). A 90x gap that is all engineering, with not a single weight changed — this is why inference engineering exists as a discipline.
An honest counterpoint
The gains of inference optimization are conditional: quantization degrades accuracy on sensitive models (especially INT4 weight-only, which is sensitive to activation outliers); speculative decoding gains little when compute is already saturated (it is limited by HBM bandwidth); PagedAttention helps less at batch=1 than at large batches. Every optimization has its sweet spot — quoting speedups without the workload is meaningless. The systematic list of traps is in Common Pitfalls and Anti-Patterns.
6. How to Use This Handbook
Inference acceleration is the intersection of engineering, algorithms, and hardware. This site is organized in the order "build concepts → understand mechanisms → dissect cases → get hands-on":
- Guides (you are here): next read Inference vs. Training vs. Fine-Tuning to draw the boundaries among the three workloads, use A Brief History to build a timeline, then Anatomy of the Overall Architecture for the full site map, and finally pick a route from Learning Paths: Three Routes.
- Core knowledge: lay the performance-analysis foundation with Latency, Throughput, and Concurrency → The GPU Memory Hierarchy and the Bandwidth Wall → The Roofline Model and Compute Analysis; then Model Quantization Fundamentals, Pruning and Sparsification, Knowledge Distillation for the model layer; Kernel Fusion and Custom Kernels, Computation Graph Optimization for the operator/graph layers; GPU Architecture and Optimization, Batching and Request Scheduling, Model Serving and Orchestration for the system/hardware layer.
- Case studies: land the abstract concepts on concrete engines — TensorRT and GPU Inference, ONNX Runtime: Cross-Platform, OpenVINO and CPU Inference, vLLM and PagedAttention, TensorRT-LLM, Speculative Decoding and Medusa/EAGLE, llama.cpp and GGUF, Mobile Deployment, Triton Inference Server, Distributed Inference (TP/PP).
- Hands-on practice: Deploy an Inference Service from Scratch, Progressive Tutorial: Three Working Versions, Inference Benchmarking in Practice, Portfolio Projects, Inference Engine Comparison, Tuning and Performance Optimization, Deployment Design Principles, Common Pitfalls and Anti-Patterns.
- Look up anytime: Glossary, Hardware Primer, Benchmark Data & Tool Profiles, Curated Resources.
Further Reading
- Inference vs. Training vs. Fine-Tuning — the full expansion of the boundary analysis in Section 2
- Anatomy of the Overall Architecture — a five-layer map of the inference system
- A Brief History — seventy years from CPU inference to the LLM inference revolution
- Latency, Throughput, and Concurrency — precise definitions of inference performance metrics
- The GPU Memory Hierarchy and the Bandwidth Wall — why LLM inference is memory-bound
- Model Quantization Fundamentals — the core technique of model-layer optimization
- vLLM and PagedAttention — the benchmark case of system-layer optimization
- Deploy an Inference Service from Scratch — grow this page's snippets into a complete demo
References
- Kwon et al. Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023) — the original vLLM and PagedAttention paper
- Dao et al. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (NeurIPS 2022) — the benchmark for operator-level optimization
- Frantar et al. GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers (ICLR 2023) — the landmark of INT4 weight quantization
- Lin et al. AWQ: Activation-Aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024) — activation-aware weight quantization
- Leviathan et al. Fast Inference from Transformers via Speculative Decoding (ICML 2023) — the early engineering of speculative decoding
- Li et al. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty (ICML 2024) — the start of the EAGLE series
- NVIDIA. TensorRT Developer Guide — official documentation of TensorRT graph and operator optimization
- NVIDIA. TensorRT-LLM — the open-source repo and benchmarks of the LLM inference engine