Appearance
Progressive Tutorial: Three Working Versions
Adding one optimization at a time is the fastest path to learning inference optimization — every step shows a change in the numbers, lets you pinpoint where the gain comes from, and can be rolled back for comparison. This article breaks "deploy an inference service from scratch" into six versions, changing exactly one thing per version.
Deploy an Inference Service from Scratch presents the "big-version" route of v1 → v2 → v3 — each jump introduces a whole bundle of optimizations, the performance leaps are dramatic, but so are the blind spots: you see v2 running 30× faster than v1, yet you can't tell how much of that 30× comes from PagedAttention, how much from continuous batching, and how much from FlashAttention. This article takes the same model (Llama-2-7B-Chat) and splits the optimization path finer: each version adds exactly one thing, with diff-style code and benchmark data, so you can see the real value of every single optimization.
The map of the six versions:
v0 baseline HuggingFace transformers + bf16 + naive generate
│
▼
v1 + KV cache Turn on use_cache=True (it's on by default;
this version explains it properly)
│
▼
v2 + continuous batching vLLM replaces transformers, introducing PagedAttention
│
▼
v3 + INT4 quantization AWQ quantization, weights 14GB → 4GB
│
▼
v4 + speculative decoding EAGLE-3 speculative decoding, latency halved at small batch
│
▼
v5 + Triton + autoscale Unified orchestration layer, horizontal scaling,
monitoring and alertsEach version is benchmarked against the previous one — this is the "control variables" method from Inference Benchmarking in Practice.
How this article relates to build-your-own
- Want the complete end-to-end project, directory structure, and go-live process? Read Deploy an Inference Service from Scratch
- Want to understand each optimization's contribution and configuration one by one? Read this article
- Both use the same model, the same hardware, and the same test method, so the numbers cross-reference cleanly
1. v0: Baseline
v0's goal is to establish the control baseline with "no optimization applied." Every later gain is measured from here.
1. Code
python
# v0_baseline.py
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(
MODEL,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
def generate(prompt, max_new_tokens=256):
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.inference_mode():
out = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=False, # greedy, for reproducibility
pad_token_id=tokenizer.eos_token_id,
)
return tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(generate("Explain PagedAttention in three sentences."))2. Benchmark
A100 40GB, single request, prompt 50 tokens / output 256 tokens:
| Metric | v0 value |
|---|---|
| TTFT (ms) | 850 |
| TPOT (ms/tok) | 40 |
| End-to-end (s) | 11.1 |
| Peak memory (GB) | 14.6 |
| Throughput (tok/s) | 25 |
That is the baseline. Every version below is compared against it.
What v0 deliberately does not use
v0 deliberately avoids vLLM, FlashAttention, and PagedAttention. Hugging Face transformers defaults to use_cache=True, so v0 actually already has KV cache reuse — keep this in mind, or v1's "turn on KV cache" will confuse you: it's on by default, so what is there to "turn on"?
Answer: v1 is really about how use_cache differs across implementations — transformers' KV cache is a list of tensors, per-layer and non-contiguous in layout; vLLM's is PagedAttention's paged layout. That is v1's real question.
2. v1: KV Cache Optimization and FlashAttention
v1's goal: swap the KV cache implementation for the "predecessors" of FlashAttention2 + PagedAttention — the optimized attention kernels of flash_attn and xformers. This version does not adopt vLLM; it only swaps the attention kernel.
1. Code Diff
python
# v1_kv_flash.py
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(
MODEL,
torch_dtype=torch.bfloat16,
device_map="auto",
attn_implementation="flash_attention_2", # ← the core change in v1
)
model.eval()attn_implementation="flash_attention_2" is an interface introduced in transformers 4.36+, switching attention computation from PyTorch's naive implementation to the FlashAttention2 kernel. See Kernel Fusion and Custom Kernels.
2. Benchmark
| Metric | v0 | v1 | v1 vs. v0 |
|---|---|---|---|
| TTFT (ms) | 850 | 600 | −29% |
| TPOT (ms/tok) | 40 | 28 | −30% |
| End-to-end (s) | 11.1 | 7.8 | −30% |
| Peak memory (GB) | 14.6 | 13.5 | −8% |
Swapping the attention kernel alone buys a 30% speedup — this is the memory-bound nature of LLM inference in action: switch to a kernel with less memory traffic and performance jumps immediately. See The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis.
Why memory usage dropped too
FlashAttention2 "dissolves" the attention intermediate matrices ([batch, head, seq, seq]) from HBM into SRAM, so it no longer needs O(seq²) of temporary space in HBM — and the peak memory drops with it. This is the canonical example of "fusion saves memory traffic" in Kernel Fusion and Custom Kernels.
3. v2: Continuous Batching (vLLM Swap)
v1 is still inside the transformers framework — the generate interface is essentially "batch=1 serial processing." v2 replaces the entire inference stack with vLLM, introducing PagedAttention and continuous batching.
1. Code Diff
bash
# v2_vllm.sh
vllm serve meta-llama/Llama-2-7b-chat-hf \
--dtype bfloat16 \
--max-model-len 4096 \
--max-num-seqs 256 \
--gpu-memory-utilization 0.9 \
--port 8000The client changes from direct Python calls to HTTP:
python
# v2_client.py
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
prompt="Explain PagedAttention in three sentences.",
max_tokens=256,
temperature=0,
)
print(resp.choices[0].text)2. Benchmark
| Metric | v1 | v2 single request | v2 32 concurrency |
|---|---|---|---|
| TTFT (ms) | 600 | 200 | 250 |
| TPOT (ms/tok) | 28 | 20 | 22 |
| End-to-end (s) | 7.8 | 5.3 | 6.0 |
| Peak memory (GB) | 13.5 | 18 | 36 |
| Throughput (tok/s) | 36 | 50 | 1100 |
The single-request gain is modest (1.4×) — PagedAttention's advantage is small in the single-request case. The real explosion is under concurrency: at 32 concurrent requests, throughput hits 1100 tok/s — 30× a single v1 request. This is the core payoff of Batching and Request Scheduling.
Why memory jumped so much
v2 memory grew from 13.5 GB to 36 GB — because vLLM carves 90% of GPU memory into a KV pool, ready to accept more concurrency. In the single-request test, most of that 18 GB KV pool sits idle, but with it, new requests can be admitted with zero latency when concurrency arrives. This is the throughput-oriented design philosophy, completely different from v1's latency-oriented one. See the "latency-first vs throughput-first" section of Deployment Design Principles.
4. v3: INT4 Quantization
v2 already saturated throughput, but memory is tight — on a single A100 40GB, the KV pool for a 7B model takes 22 GB, and the concurrency ceiling is capped by memory. v3 uses AWQ INT4 quantization to compress weights from 14 GB to 4 GB, and hands all 10 GB freed to the KV cache so it can accept more concurrency.
1. Code Diff
bash
# v3_awq.sh
vllm serve ./Llama-2-7b-chat-hf-awq \
--quantization awq \
--dtype float16 \
--max-model-len 4096 \
--max-num-seqs 384 \ # ← raise the concurrency ceiling since memory freed up
--gpu-memory-utilization 0.9 \
--port 80002. Benchmark
| Metric | v2 | v3 | v3 vs. v2 |
|---|---|---|---|
| Weight memory (GB) | 14 | 4 | −71% |
| Single-request TPOT (ms/tok) | 20 | 14 | −30% |
| Throughput at 32 concurrency (tok/s) | 1100 | 1700 | +55% |
| Throughput at 64 concurrency (tok/s) | 720 (queuing begins) | 2100 | +190% |
| Model quality (MMLU) | 45.3 | 44.8 | −0.5 |
v3's gain at low concurrency is unremarkable (30%); at high concurrency the advantage multiplies — quantization reduces memory traffic, and the larger KV pool admits more concurrency. This is the classic signature of a memory-bound workload.
Why TPOT dropped too
At 32 concurrency, v2's TPOT was actually higher than single-request (22 vs 20) — because KV cache memory ran short and some requests had their KV swapped out or recomputed. After v3's quantization frees up memory, the KV pool is ample and TPOT comes down naturally. This is the textbook case of "balancing gpu_memory_utilization and max_num_seqs" in Tuning and Performance Optimization.
5. v4: Speculative Decoding
v3's throughput is already strong, but single-request latency is stuck at 14 ms/tok. For scenarios where "one person is waiting for the result" (chat, code completion), throughput doesn't matter — TPOT is the lifeline. v4 introduces EAGLE-3 speculative decoding: the small model guesses tokens, the large model verifies them in one pass, and TPOT is cut in half.
1. Code Diff
bash
# v4_speculative.sh
vllm serve ./Llama-2-7b-chat-hf-awq \
--quantization awq \
--speculative-model "lmsys/Falcon-EAGLE-3-7B" \
--num-speculative-tokens 5 \
--speculative-draft-tensor-parallel-size 1 \
--max-model-len 4096 \
--max-num-seqs 16 \ # ← deliberately small, see explanation below
--port 80002. Benchmark
| Metric | v3 | v4 single request | v4 16 concurrency | v4 64 concurrency |
|---|---|---|---|---|
| Single-request TPOT (ms/tok) | 14 | 6 | 7 | 12 |
| Single-request end-to-end (s) | 3.6 | 1.6 | 1.8 | 3.1 |
| Throughput at 16 concurrency (tok/s) | 700 | 700 | 2200 | - |
| Throughput at 64 concurrency (tok/s) | 2100 | - | - | 1800 (regression) |
| Peak memory (GB) | 22 | 36 (extra draft model) | 36 | 36 |
v4 cuts single-request TPOT in half at small batch (14 → 6), with throughput doubling in step. But at large batch (64 concurrency), performance actually regresses — because the large model is already compute-saturated at large batch, and speculative decoding only adds memory and compute. This is the core trade-off of Speculative Decoding and Medusa/EAGLE.
The golden range of speculative decoding
v4 pays off most at concurrency ≤ 16; 32 is the break-even point; beyond 64 it is a net loss. In production, either route low-concurrency traffic to v4 with a router, or protect v4 with a dynamic batch ceiling — never let 64-concurrency traffic flood a v4 instance. See Tuning and Performance Optimization and Common Pitfalls and Anti-Patterns.
6. v5: Triton + Autoscale
v0–v4 all make "a single instance run fast," but production still has a last mile: horizontal scaling, multi-model coexistence, gray releases, monitoring and alerts. v5 uses Triton Inference Server as the unified orchestration layer, plus K8s HPA for autoscaling.
1. Architecture Diff
┌──────────────────────────────────┐
│ Load Balancer (LB) │
└──────────────┬───────────────────┘
▼
┌──────────────────────────────────┐
│ K8s HPA (scale on QPS/latency) │
└──────────────┬───────────────────┘
▼
┌─────────────┬─────────────┼─────────────┬──────────────┐
▼ ▼ ▼ ▼ ▼
Triton-1 Triton-2 Triton-3 Triton-4 Triton-N
(vLLM+AWQ (vLLM+AWQ (vLLM+AWQ (vLLM+AWQ (...)
+EAGLE) +EAGLE) +EAGLE) +EAGLE)
│ │ │ │
└─────────────┴─────────────┴─────────────┘
│
▼
/metrics (Prometheus)2. Config Diff
model_repository/llama2-7b-awq/config.pbtxt:
name: "llama2-7b-awq"
backend: "vllm" # vLLM 0.6+ native Triton backend
max_batch_size: 16
dynamic_batching {
preferred_batch_size: [ 4, 8, 16 ]
max_queue_delay_microseconds: 50000
}
instance_group [
{ kind: KIND_AUTO, count: 1, name: "vllm" }
]
parameters {
key: "model"
value: { string_value: "/models/llama2-7b-chat-hf-awq" }
}
parameters {
key: "quantization"
value: { string_value: "awq" }
}K8s HPA (scaling on QPS and p99 latency):
yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: triton-vllm-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: triton-vllm
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: vllm_requests_per_second
target:
type: AverageValue
averageValue: "20" # scale out at 20 QPS per instance
- type: Pods
pods:
metric:
name: vllm_p99_latency_ms
target:
type: AverageValue
averageValue: "3000" # scale out when p99 exceeds 3s3. Benchmark
| Metric | v4 single instance | v5 three instances | v5 ten instances |
|---|---|---|---|
| Per-instance QPS ceiling | 20 | 20 | 20 |
| Cluster QPS ceiling | 20 | 60 | 200 |
| Cluster p99 latency (s) | 1.8 | 1.8 | 2.1 (slightly up) |
| Impact of a single-instance failure | Whole site down | Some requests 5xx | Imperceptible (traffic retries) |
v5's gain is not "a single instance runs faster" but horizontal scalability and failure recovery. This is the qualitative leap from "a fast model" to "an operable service." See Model Serving and Orchestration and Triton Inference Server.
How to set HPA thresholds
- Per-instance QPS ceiling: measure it once with a load test (20 in this article); set the HPA trigger at 70% of the ceiling (scale out at 14 QPS)
- p99 latency threshold: 60% of the business SLA (SLA 5s → HPA scales out at 3s)
- Scale cooldown: the default 5 minutes is too long for inference services; use 1–2 minutes — request peaks arrive and leave fast
See the "autoscale tuning" section of Tuning and Performance Optimization.
7. Full Picture of the Six Versions
Here is the consolidated metrics table (A100 40GB, Llama-2-7B-Chat, single request prompt 50 / output 256, concurrency 32):
| Version | Key optimization | TTFT (ms) | TPOT (ms) | Throughput at 32 concurrency (tok/s) | Memory (GB) | Model quality |
|---|---|---|---|---|---|---|
| v0 | baseline | 850 | 40 | ≈ 25 | 14.6 | 45.3 |
| v1 | +FlashAttention2 | 600 | 28 | ≈ 36 | 13.5 | 45.3 |
| v2 | +PagedAttention + continuous batching | 200 | 20 | 1100 | 36 | 45.3 |
| v3 | +AWQ INT4 | 180 | 14 | 1700 | 22 | 44.8 |
| v4 | +EAGLE speculative decoding | 150 | 6 | 2200 (at 16 concurrency) | 36 | 44.8 |
| v5 | +Triton + autoscale | 200 | 22 | Cluster-level horizontal scaling | 22×N | 44.8 |
The incremental gain of each version over the previous one:
v0 → v1: FlashAttention2 +30% (swap a kernel, smallest cost)
v1 → v2: PagedAttention+CB +30× (architectural leap: single batch → continuous batch)
v2 → v3: AWQ INT4 +55% (quantization compression: memory for concurrency)
v3 → v4: EAGLE speculative +95% (works only at small batch, regresses at large batch)
v4 → v5: Triton + autoscale no single-instance gain, but QPS scales horizontally to 10×8. Suggested Order of Hands-On Work
Running v0 through v5 in this article's order is the most reliable learning path:
- v0 + v1: half a day, mastering the transformers + FlashAttention basics
- v2: one day on vLLM, understanding PagedAttention and continuous batching — the core of Batching and Request Scheduling
- v3: half a day on AWQ, understanding the trade-offs in Model Quantization Fundamentals
- v4: half a day on EAGLE, understanding the batch dependency of speculative decoding
- v5: two or three days on Triton + K8s, understanding Model Serving and Orchestration
The anti-pattern of jumping from v1 straight to v3
Some teams go straight to quantization + speculative decoding, skipping vLLM — and end up with miscalculated memory budgets, misconfigured KV pools, and failing performance. v2's PagedAttention is the prerequisite foundation for v3/v4: without unified KV cache management, the memory saved by quantization can't be routed into the KV pool either. Go in order, and every step's gain will reproduce.
9. Further Reading
- Deploy an Inference Service from Scratch — the "big-version" counterpart of this article, one bundle of optimizations per version
- Inference Benchmarking in Practice — the methodology behind this article's benchmark data
- Tuning and Performance Optimization — the complete tuning checklist for every parameter in v2–v5
- Deployment Design Principles — 12 disciplines across versions
- Common Pitfalls and Anti-Patterns — pitfalls you can hit in every version
- vLLM and PagedAttention — deep dive into v2's engine
- Speculative Decoding and Medusa/EAGLE — deep dive into v4's engine
- Triton Inference Server — deep dive into v5's orchestration layer
- Model Quantization Fundamentals — the principles behind v3
- Batching and Request Scheduling — the core mechanism of v2
References
- vLLM Documentation: Quickstart — v2 launch CLI
- vLLM Documentation: Quantization — v3 AWQ configuration
- vLLM Documentation: Speculative Decoding — v4 EAGLE configuration
- HuggingFace Transformers: FlashAttention2 — the v1 attn_implementation parameter
- AutoAWQ: Quickstart — v3 quantization script
- Triton Inference Server: vLLM Backend — v5 Triton configuration
- Kubernetes HPA Documentation — v5 autoscale
- EAGLE-3: Speculative Decoding — v4 draft model