Skip to content

Progressive Tutorial: Three Working Versions

At a glance Break the v1→v2→v3 deployment optimization into six versions (v0 baseline → v1 KV cache → v2 continuous batching → v3 INT4 → v4 speculative decoding → v5 Triton + autoscale), adding exactly one optimization per version, with diff-style code and the corresponding benchmark data so you can learn them one by one.

Progressive Tutorial: Three Working Versions ​

Adding one optimization at a time is the fastest path to learning inference optimization — every step shows a change in the numbers, lets you pinpoint where the gain comes from, and can be rolled back for comparison. This article breaks "deploy an inference service from scratch" into six versions, changing exactly one thing per version.

Deploy an Inference Service from Scratch presents the "big-version" route of v1 → v2 → v3 — each jump introduces a whole bundle of optimizations, the performance leaps are dramatic, but so are the blind spots: you see v2 running 30× faster than v1, yet you can't tell how much of that 30× comes from PagedAttention, how much from continuous batching, and how much from FlashAttention. This article takes the same model (Llama-2-7B-Chat) and splits the optimization path finer: each version adds exactly one thing, with diff-style code and benchmark data, so you can see the real value of every single optimization.

The map of the six versions:

v0  baseline               HuggingFace transformers + bf16 + naive generate
     │
     ▼
v1  + KV cache             Turn on use_cache=True (it's on by default;
                            this version explains it properly)
     │
     ▼
v2  + continuous batching  vLLM replaces transformers, introducing PagedAttention
     │
     ▼
v3  + INT4 quantization    AWQ quantization, weights 14GB → 4GB
     │
     ▼
v4  + speculative decoding EAGLE-3 speculative decoding, latency halved at small batch
     │
     ▼
v5  + Triton + autoscale   Unified orchestration layer, horizontal scaling,
                            monitoring and alerts

Each version is benchmarked against the previous one — this is the "control variables" method from Inference Benchmarking in Practice.

How this article relates to build-your-own

  • Want the complete end-to-end project, directory structure, and go-live process? Read Deploy an Inference Service from Scratch
  • Want to understand each optimization's contribution and configuration one by one? Read this article
  • Both use the same model, the same hardware, and the same test method, so the numbers cross-reference cleanly

1. v0: Baseline ​

v0's goal is to establish the control baseline with "no optimization applied." Every later gain is measured from here.

1. Code ​

python
# v0_baseline.py
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "meta-llama/Llama-2-7b-chat-hf"

tokenizer = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(
    MODEL,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
model.eval()

def generate(prompt, max_new_tokens=256):
    inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
    with torch.inference_mode():
        out = model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            do_sample=False,            # greedy, for reproducibility
            pad_token_id=tokenizer.eos_token_id,
        )
    return tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)

print(generate("Explain PagedAttention in three sentences."))

2. Benchmark ​

A100 40GB, single request, prompt 50 tokens / output 256 tokens:

Metricv0 value
TTFT (ms)850
TPOT (ms/tok)40
End-to-end (s)11.1
Peak memory (GB)14.6
Throughput (tok/s)25

That is the baseline. Every version below is compared against it.

What v0 deliberately does not use

v0 deliberately avoids vLLM, FlashAttention, and PagedAttention. Hugging Face transformers defaults to use_cache=True, so v0 actually already has KV cache reuse — keep this in mind, or v1's "turn on KV cache" will confuse you: it's on by default, so what is there to "turn on"?

Answer: v1 is really about how use_cache differs across implementations — transformers' KV cache is a list of tensors, per-layer and non-contiguous in layout; vLLM's is PagedAttention's paged layout. That is v1's real question.

2. v1: KV Cache Optimization and FlashAttention ​

v1's goal: swap the KV cache implementation for the "predecessors" of FlashAttention2 + PagedAttention — the optimized attention kernels of flash_attn and xformers. This version does not adopt vLLM; it only swaps the attention kernel.

1. Code Diff ​

python
# v1_kv_flash.py
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "meta-llama/Llama-2-7b-chat-hf"

tokenizer = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(
    MODEL,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",    # ← the core change in v1
)
model.eval()

attn_implementation="flash_attention_2" is an interface introduced in transformers 4.36+, switching attention computation from PyTorch's naive implementation to the FlashAttention2 kernel. See Kernel Fusion and Custom Kernels.

2. Benchmark ​

Metricv0v1v1 vs. v0
TTFT (ms)850600−29%
TPOT (ms/tok)4028−30%
End-to-end (s)11.17.8−30%
Peak memory (GB)14.613.5−8%

Swapping the attention kernel alone buys a 30% speedup — this is the memory-bound nature of LLM inference in action: switch to a kernel with less memory traffic and performance jumps immediately. See The GPU Memory Hierarchy and the Bandwidth Wall and The Roofline Model and Compute Analysis.

Why memory usage dropped too

FlashAttention2 "dissolves" the attention intermediate matrices ([batch, head, seq, seq]) from HBM into SRAM, so it no longer needs O(seq²) of temporary space in HBM — and the peak memory drops with it. This is the canonical example of "fusion saves memory traffic" in Kernel Fusion and Custom Kernels.

3. v2: Continuous Batching (vLLM Swap) ​

v1 is still inside the transformers framework — the generate interface is essentially "batch=1 serial processing." v2 replaces the entire inference stack with vLLM, introducing PagedAttention and continuous batching.

1. Code Diff ​

bash
# v2_vllm.sh
vllm serve meta-llama/Llama-2-7b-chat-hf \
    --dtype bfloat16 \
    --max-model-len 4096 \
    --max-num-seqs 256 \
    --gpu-memory-utilization 0.9 \
    --port 8000

The client changes from direct Python calls to HTTP:

python
# v2_client.py
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.completions.create(
    model="meta-llama/Llama-2-7b-chat-hf",
    prompt="Explain PagedAttention in three sentences.",
    max_tokens=256,
    temperature=0,
)
print(resp.choices[0].text)

2. Benchmark ​

Metricv1v2 single requestv2 32 concurrency
TTFT (ms)600200250
TPOT (ms/tok)282022
End-to-end (s)7.85.36.0
Peak memory (GB)13.51836
Throughput (tok/s)36501100

The single-request gain is modest (1.4×) — PagedAttention's advantage is small in the single-request case. The real explosion is under concurrency: at 32 concurrent requests, throughput hits 1100 tok/s — 30× a single v1 request. This is the core payoff of Batching and Request Scheduling.

Why memory jumped so much

v2 memory grew from 13.5 GB to 36 GB — because vLLM carves 90% of GPU memory into a KV pool, ready to accept more concurrency. In the single-request test, most of that 18 GB KV pool sits idle, but with it, new requests can be admitted with zero latency when concurrency arrives. This is the throughput-oriented design philosophy, completely different from v1's latency-oriented one. See the "latency-first vs throughput-first" section of Deployment Design Principles.

4. v3: INT4 Quantization ​

v2 already saturated throughput, but memory is tight — on a single A100 40GB, the KV pool for a 7B model takes 22 GB, and the concurrency ceiling is capped by memory. v3 uses AWQ INT4 quantization to compress weights from 14 GB to 4 GB, and hands all 10 GB freed to the KV cache so it can accept more concurrency.

1. Code Diff ​

bash
# v3_awq.sh
vllm serve ./Llama-2-7b-chat-hf-awq \
    --quantization awq \
    --dtype float16 \
    --max-model-len 4096 \
    --max-num-seqs 384 \           # ← raise the concurrency ceiling since memory freed up
    --gpu-memory-utilization 0.9 \
    --port 8000

2. Benchmark ​

Metricv2v3v3 vs. v2
Weight memory (GB)144−71%
Single-request TPOT (ms/tok)2014−30%
Throughput at 32 concurrency (tok/s)11001700+55%
Throughput at 64 concurrency (tok/s)720 (queuing begins)2100+190%
Model quality (MMLU)45.344.8−0.5

v3's gain at low concurrency is unremarkable (30%); at high concurrency the advantage multiplies — quantization reduces memory traffic, and the larger KV pool admits more concurrency. This is the classic signature of a memory-bound workload.

Why TPOT dropped too

At 32 concurrency, v2's TPOT was actually higher than single-request (22 vs 20) — because KV cache memory ran short and some requests had their KV swapped out or recomputed. After v3's quantization frees up memory, the KV pool is ample and TPOT comes down naturally. This is the textbook case of "balancing gpu_memory_utilization and max_num_seqs" in Tuning and Performance Optimization.

5. v4: Speculative Decoding ​

v3's throughput is already strong, but single-request latency is stuck at 14 ms/tok. For scenarios where "one person is waiting for the result" (chat, code completion), throughput doesn't matter — TPOT is the lifeline. v4 introduces EAGLE-3 speculative decoding: the small model guesses tokens, the large model verifies them in one pass, and TPOT is cut in half.

1. Code Diff ​

bash
# v4_speculative.sh
vllm serve ./Llama-2-7b-chat-hf-awq \
    --quantization awq \
    --speculative-model "lmsys/Falcon-EAGLE-3-7B" \
    --num-speculative-tokens 5 \
    --speculative-draft-tensor-parallel-size 1 \
    --max-model-len 4096 \
    --max-num-seqs 16 \            # ← deliberately small, see explanation below
    --port 8000

2. Benchmark ​

Metricv3v4 single requestv4 16 concurrencyv4 64 concurrency
Single-request TPOT (ms/tok)146712
Single-request end-to-end (s)3.61.61.83.1
Throughput at 16 concurrency (tok/s)7007002200-
Throughput at 64 concurrency (tok/s)2100--1800 (regression)
Peak memory (GB)2236 (extra draft model)3636

v4 cuts single-request TPOT in half at small batch (14 → 6), with throughput doubling in step. But at large batch (64 concurrency), performance actually regresses — because the large model is already compute-saturated at large batch, and speculative decoding only adds memory and compute. This is the core trade-off of Speculative Decoding and Medusa/EAGLE.

The golden range of speculative decoding

v4 pays off most at concurrency ≤ 16; 32 is the break-even point; beyond 64 it is a net loss. In production, either route low-concurrency traffic to v4 with a router, or protect v4 with a dynamic batch ceiling — never let 64-concurrency traffic flood a v4 instance. See Tuning and Performance Optimization and Common Pitfalls and Anti-Patterns.

6. v5: Triton + Autoscale ​

v0–v4 all make "a single instance run fast," but production still has a last mile: horizontal scaling, multi-model coexistence, gray releases, monitoring and alerts. v5 uses Triton Inference Server as the unified orchestration layer, plus K8s HPA for autoscaling.

1. Architecture Diff ​

                ┌──────────────────────────────────┐
                │        Load Balancer (LB)         │
                └──────────────┬───────────────────┘
                               ▼
                ┌──────────────────────────────────┐
                │  K8s HPA (scale on QPS/latency)   │
                └──────────────┬───────────────────┘
                               ▼
   ┌─────────────┬─────────────┼─────────────┬──────────────┐
   ▼             ▼             ▼             ▼              ▼
 Triton-1     Triton-2      Triton-3      Triton-4       Triton-N
 (vLLM+AWQ   (vLLM+AWQ     (vLLM+AWQ     (vLLM+AWQ      (...)
  +EAGLE)     +EAGLE)       +EAGLE)       +EAGLE)
   │             │             │             │
   └─────────────┴─────────────┴─────────────┘
                               │
                               ▼
                       /metrics (Prometheus)

2. Config Diff ​

model_repository/llama2-7b-awq/config.pbtxt:

name: "llama2-7b-awq"
backend: "vllm"               # vLLM 0.6+ native Triton backend
max_batch_size: 16
dynamic_batching {
  preferred_batch_size: [ 4, 8, 16 ]
  max_queue_delay_microseconds: 50000
}
instance_group [
  { kind: KIND_AUTO, count: 1, name: "vllm" }
]
parameters {
  key: "model"
  value: { string_value: "/models/llama2-7b-chat-hf-awq" }
}
parameters {
  key: "quantization"
  value: { string_value: "awq" }
}

K8s HPA (scaling on QPS and p99 latency):

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: triton-vllm-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: triton-vllm
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Pods
      pods:
        metric:
          name: vllm_requests_per_second
        target:
          type: AverageValue
          averageValue: "20"      # scale out at 20 QPS per instance
    - type: Pods
      pods:
        metric:
          name: vllm_p99_latency_ms
        target:
          type: AverageValue
          averageValue: "3000"    # scale out when p99 exceeds 3s

3. Benchmark ​

Metricv4 single instancev5 three instancesv5 ten instances
Per-instance QPS ceiling202020
Cluster QPS ceiling2060200
Cluster p99 latency (s)1.81.82.1 (slightly up)
Impact of a single-instance failureWhole site downSome requests 5xxImperceptible (traffic retries)

v5's gain is not "a single instance runs faster" but horizontal scalability and failure recovery. This is the qualitative leap from "a fast model" to "an operable service." See Model Serving and Orchestration and Triton Inference Server.

How to set HPA thresholds

  • Per-instance QPS ceiling: measure it once with a load test (20 in this article); set the HPA trigger at 70% of the ceiling (scale out at 14 QPS)
  • p99 latency threshold: 60% of the business SLA (SLA 5s → HPA scales out at 3s)
  • Scale cooldown: the default 5 minutes is too long for inference services; use 1–2 minutes — request peaks arrive and leave fast

See the "autoscale tuning" section of Tuning and Performance Optimization.

7. Full Picture of the Six Versions ​

Here is the consolidated metrics table (A100 40GB, Llama-2-7B-Chat, single request prompt 50 / output 256, concurrency 32):

VersionKey optimizationTTFT (ms)TPOT (ms)Throughput at 32 concurrency (tok/s)Memory (GB)Model quality
v0baseline85040≈ 2514.645.3
v1+FlashAttention260028≈ 3613.545.3
v2+PagedAttention + continuous batching2002011003645.3
v3+AWQ INT41801417002244.8
v4+EAGLE speculative decoding15062200 (at 16 concurrency)3644.8
v5+Triton + autoscale20022Cluster-level horizontal scaling22×N44.8

The incremental gain of each version over the previous one:

v0 → v1:  FlashAttention2       +30%   (swap a kernel, smallest cost)
v1 → v2:  PagedAttention+CB     +30×   (architectural leap: single batch → continuous batch)
v2 → v3:  AWQ INT4              +55%   (quantization compression: memory for concurrency)
v3 → v4:  EAGLE speculative     +95%   (works only at small batch, regresses at large batch)
v4 → v5:  Triton + autoscale    no single-instance gain, but QPS scales horizontally to 10×

8. Suggested Order of Hands-On Work ​

Running v0 through v5 in this article's order is the most reliable learning path:

  1. v0 + v1: half a day, mastering the transformers + FlashAttention basics
  2. v2: one day on vLLM, understanding PagedAttention and continuous batching — the core of Batching and Request Scheduling
  3. v3: half a day on AWQ, understanding the trade-offs in Model Quantization Fundamentals
  4. v4: half a day on EAGLE, understanding the batch dependency of speculative decoding
  5. v5: two or three days on Triton + K8s, understanding Model Serving and Orchestration

The anti-pattern of jumping from v1 straight to v3

Some teams go straight to quantization + speculative decoding, skipping vLLM — and end up with miscalculated memory budgets, misconfigured KV pools, and failing performance. v2's PagedAttention is the prerequisite foundation for v3/v4: without unified KV cache management, the memory saved by quantization can't be routed into the KV pool either. Go in order, and every step's gain will reproduce.

9. Further Reading ​

References ​