Appearance
Deploy an Inference Service from Scratch
Reading ten benchmark reports is worth less than carrying one deployment through end to end. This article walks you through the full journey of an inference service with the classic Llama-2-7B — from "it runs" to "it runs fast" to "it runs steady." This is not about calling a library and getting a number: it explains why each version is built the way it is, where the gains come from, and what the trade-offs cost.
Many beginners, after learning a pile of inference optimization techniques, still don't know what a production-grade inference service actually looks like. How do you lift v1 straight out of the training framework? Does swapping vLLM in for transformers really buy you several times the performance? Will quantization and speculative decoding fight each other when stacked? This article runs the complete deployment flow from start to finish along a single throughline: a Llama-2-7B chat service. What you walk away with is not "this model's tokens/s number" but an optimization skeleton you can reuse for any LLM deployment.
The full project ships in three versions (v1 → v2 → v3). Here is the map first:
Starting point: Llama-2-7B weights on Hugging Face
│
▼
┌─────────────────────────────────────────┐
│ v1: PyTorch + Transformers, direct │
│ - transformers from_pregressive │
│ - bf16, single batch, no KV cache │
│ reuse │
│ - Observe latency / memory baseline │
└──────────────────┬──────────────────────┘
▼
┌─────────────────────────────────────────┐
│ v2: vLLM replaces transformers │
│ - PagedAttention + continuous batching │
│ - OpenAI-compatible API service │
│ - 5–10× throughput, p99 latency down │
└──────────────────┬──────────────────────┘
▼
┌─────────────────────────────────────────┐
│ v3: AWQ INT4 + EAGLE-3 + Triton │
│ - 4-bit quantization compresses weights │
│ - Speculative decoding for extreme │
│ latency at small batch │
│ - Triton orchestrates multiple models │
│ and versions │
└─────────────────────────────────────────┘The three versions answer three typical engineering questions: v1 answers "can it run at all," v2 answers "is throughput good enough," and v3 answers "can cost and latency be cut further." They are exactly the standard upgrade path for inference optimization in industry — see Learning Paths: Three Routes and Batching and Request Scheduling.
Prerequisites
This article assumes you already understand the basics of LLM inference: Latency, Throughput, and Concurrency, The GPU Memory Hierarchy and the Bandwidth Wall, Model Quantization Fundamentals, and Model Serving and Orchestration. If you still need to catch up, read What Is Inference Acceleration? and Anatomy of the Overall Architecture first, then come back.
1. Choosing the Project: Why Llama-2-7B
The first question: which model should you pick for a hands-on project? There are three criteria, and all three are mandatory.
| Criterion | Why | Cost of violating it |
|---|---|---|
| Fits on a single GPU (~≤ 14B params) | One A100/L40/4090 runs the whole flow without tensor-parallel headaches | Multi-GPU distributed inflates the complexity 10× — you lose the main thread |
| Broad engine support (vLLM/TRT-LLM/llama.cpp all support it) | Enables cross-engine comparison — comparison is the shortcut to learning | Pick an obscure model with no engine support and half the optimization techniques can't be verified |
| A public chat version available | Lets you run a real workload (a Q&A service), not just generate demos | With only a base model, output quality is poor and there is no end-to-end experience |
Against these criteria, the three candidates:
| Model | Params | Single-GPU memory (bf16 + KV) | Ecosystem support | Fit |
|---|---|---|---|---|
| Llama-2-7B-Chat | 7B | ≈ 16 GB | ★★★ Full coverage: vLLM/TRT-LLM/llama.cpp/Ollama | ★★★ First choice |
| Qwen2.5-7B-Instruct | 7B | ≈ 16 GB | ★★★ The most complete ecosystem among Chinese models | ★★★ First choice for Chinese scenarios |
| Mistral-7B-Instruct-v0.3 | 7B | ≈ 16 GB | ★★☆ Supported by vLLM/TRT-LLM | ★★☆ English scenarios |
This article picks Llama-2-7B-Chat for three reasons:
- Ecosystem benchmark: vLLM, TensorRT-LLM, and llama.cpp — the three mainstream engines — all treat it as a first-class citizen, so comparisons are fair.
- Manageable memory: bf16 inference takes about 14 GB of weights + 2 GB of KV cache, so a single A100 (40 GB) or RTX 4090 (24 GB) can carry all three versions of the experiment.
- Publicly downloadable: Meta's official weights are public on Hugging Face (gated access required) — no special permission needed.
Define "success" before writing any code
When you translate a problem into a deployment task, you must answer three questions:
- Workload type: online chat (streaming generation, low latency) or offline batch (high throughput)? This article takes online chat as the main thread because it is the most typical and the most demanding.
- SLA definition: time to first token (TTFT) < 500 ms, per-request generation latency (TPOT) < 50 ms/token, p99 end-to-end < 4 s, and > 20 QPS on a single GPU.
- Baseline: v1's raw PyTorch inference is the baseline. Every later optimization must clearly beat it, or the optimization is pointless.
The SLA is written down before the code. That is basic hygiene from Deployment Design Principles.
2. Setting Up the Environment
1. Version Requirements
The code in this article is based on the following versions (stable releases from 2024–2025):
| Software | Version | Purpose |
|---|---|---|
| Python | 3.10+ | The language itself |
| torch | 2.3+ / CUDA 12.1 | Training framework and GPU backend |
| transformers | 4.44+ | Hugging Face model loading |
| vllm | 0.6+ | Primary inference engine for v2/v3 |
| autoawq | 0.2.6+ | AWQ 4-bit quantization |
| tritonserver | 24.08+ | Triton inference server |
| nvidia-smi | driver ≥ 535 | GPU monitoring |
2. Create a Virtual Environment
Never install dependencies into the system Python. Inference optimization is extremely version-sensitive — a one-digit mismatch in CUDA/cuDNN/torch versions can cost you 30% of performance:
bash
mkdir llama-deploy && cd llama-deploy
python -m venv .venv
source .venv/bin/activate # macOS / Linux
# .venv\Scripts\activate # Windows PowerShell
# v1 uses: transformers + torch
pip install torch==2.3.1 transformers==4.44.2 accelerate sentencepiece
# v2 uses: vLLM (bundles PagedAttention)
pip install vllm==0.6.3
# v3 uses: AWQ quantization + Triton client
pip install autoawq==0.2.6 tritonclient[all]==2.44.0
pip freeze > requirements.txtWhy vLLM gets its own install
vLLM's wheels ship precompiled CUDA kernels (PagedAttention, FlashAttention) that demand an exact match between CUDA and torch versions. Mixed into the transformers environment, they conflict constantly. In production practice, v1 and v2 are usually split into two separate virtual environments or two separate Docker images — which is exactly why Triton later became the "unified orchestration layer." See Triton Inference Server.
3. GPU and Memory Self-Check
bash
nvidia-smi --query-gpu=name,memory.total,memory.free,driver_version,compute_cap --format=csv
# name, memory.total [MiB], memory.free [MiB], driver_version, compute_cap
# NVIDIA A100 80GB, 81920, 78000, 535.104.05, 8.0compute_cap must be ≥ 8.0 (Ampere or newer) to run this article's full bf16 + PagedAttention + FlashAttention path end to end. Older cards like the T4/V100 can still run it, but their performance numbers are in a different class.
3. v1: Raw PyTorch + Transformers Inference
v1 has exactly one goal: make the model produce decent answers. It deliberately uses no tricks; it is the control baseline for everything that follows.
1. Load the Model
python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "meta-llama/Llama-2-7b-chat-hf"
tokenizer = AutoTokenizer.from_pretrained(MODEL, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
MODEL,
torch_dtype=torch.bfloat16, # bf16 is more numerically stable than fp16 and halves memory vs fp32
device_map="auto", # automatically move weights onto the GPU
)
model.eval()
prompt = "Explain PagedAttention in three sentences."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")device_map="auto" is a capability of Hugging Face accelerate — it distributes model layers across the visible GPUs/CPU according to the memory budget. On a single GPU, it simply means "move everything to the GPU."
2. Run One Inference
python
with torch.inference_mode():
out = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
temperature=0.6,
top_p=0.9,
pad_token_id=tokenizer.eos_token_id,
)
text = tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(text)3. Observe the Baseline Metrics
The point of v1 is not "it runs" — it is measuring the baseline: single-request latency, peak memory, and tokens/s. This is the reference point for every optimization that follows.
python
import time, torch
def bench_one(prompt, max_new_tokens=256):
torch.cuda.reset_peak_memory_stats()
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
# warmup (the first inference triggers kernel compilation and must not enter the baseline)
_ = model.generate(**inputs, max_new_tokens=8, do_sample=False)
torch.cuda.synchronize()
t0 = time.perf_counter()
out = model.generate(**inputs, max_new_tokens=max_new_tokens,
do_sample=True, temperature=0.6, top_p=0.9,
pad_token_id=tokenizer.eos_token_id)
torch.cuda.synchronize()
t1 = time.perf_counter()
peak = torch.cuda.max_memory_allocated() / 1e9 # GB
n_tok = out.shape[1] - inputs["input_ids"].shape[1]
tps = n_tok / (t1 - t0)
print(f"Output {n_tok} tokens / {(t1-t0)*1000:.0f} ms = {tps:.1f} tok/s | peak {peak:.1f} GB")
return {"tokens": n_tok, "ms": (t1-t0)*1000, "tps": tps, "peak_gb": peak}
bench_one("Write a 200-word essay about autumn.")Typical v1 baseline on an A100:
| Metric | Value | Note |
|---|---|---|
| Tokens generated per request | 256 | Aligned with max_new_tokens |
| End-to-end latency | ≈ 9–11 s | Includes prefill + decode |
| Decode-phase tokens/s | ≈ 25–30 tok/s | Single batch, no PagedAttention |
| Peak memory | ≈ 14.5 GB | 7B weights 14 GB + KV cache + temporaries |
| GPU utilization | 30–50% | Decode phase is severely memory-bound |
That number is the baseline. Write it down — v2 and v3 will both be measured against it. Why are latency and GPU utilization so low? See The Roofline Model and Compute Analysis and The GPU Memory Hierarchy and the Bandwidth Wall.
The two core problems v1 exposes
- Extremely low throughput: 30 tok/s means a single GPU can serve 30 concurrent users at 1 token/s each — nowhere near a SaaS bar.
- Low GPU utilization: 30–50% means compute is sitting idle. The reason: Hugging Face's
generatedecodes token by token by default with batch=1 and a primitive KV cache reuse mechanism, so every step is limited by memory bandwidth, not compute — which is exactly the core problem Batching and Request Scheduling exists to solve.
4. v2: vLLM Replaces Transformers
v2's goal is to pull throughput up with vLLM. Its two core mechanisms — PagedAttention and continuous batching — are the biggest engineering breakthroughs in inference optimization over the past two years. See vLLM and PagedAttention.
1. Start an OpenAI-Compatible Server with the CLI
vLLM ships with the vllm serve CLI, which brings up an OpenAI-protocol HTTP server in a few lines:
bash
# Start Llama-2-7B-Chat on a single GPU with the OpenAI-compatible API
vllm serve meta-llama/Llama-2-7b-chat-hf \
--dtype bfloat16 \
--max-model-len 4096 \
--max-num-seqs 256 \
--gpu-memory-utilization 0.9 \
--enforce-eager \
--port 8000What the key parameters mean (see Tuning and Performance Optimization):
| Parameter | Meaning | v1 equivalent |
|---|---|---|
--max-model-len | Maximum context length of the model | The max_new_tokens cap in transformers |
--max-num-seqs | Maximum number of requests in flight | None in v1 — effectively batch=1 |
--gpu-memory-utilization | Fraction of total GPU memory used for the KV cache | No active management in v1 |
--enforce-eager | Disable CUDA Graphs (for debugging) | - |
2. Client Test
python
from openai import OpenAI
import time, asyncio
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
def bench_one_vllm(prompt, max_tokens=256):
t0 = time.perf_counter()
resp = client.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
prompt=prompt,
max_tokens=max_tokens,
temperature=0.6,
top_p=0.9,
)
t1 = time.perf_counter()
n_tok = resp.usage.completion_tokens
print(f"Output {n_tok} tokens / {(t1-t0)*1000:.0f} ms = {n_tok/(t1-t0):.1f} tok/s")
return {"tokens": n_tok, "ms": (t1-t0)*1000, "tps": n_tok/(t1-t0)}
bench_one_vllm("Write a 200-word essay about autumn.")3. Concurrent Load Test
v1 was slow even for a single request; vLLM's real value shows under concurrent requests. Below, 32 requests hit the server at once:
python
import asyncio
from openai import AsyncOpenAI
aclient = AsyncOpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
async def one(i):
t0 = time.perf_counter()
r = await aclient.completions.create(
model="meta-llama/Llama-2-7b-chat-hf",
prompt=f"Explain PagedAttention in three sentences. Request number {i}.",
max_tokens=128,
)
return r.usage.completion_tokens, time.perf_counter() - t0
async def main(n=32):
t0 = time.perf_counter()
results = await asyncio.gather(*[one(i) for i in range(n)])
total_toks = sum(r[0] for r in results)
total_time = time.perf_counter() - t0
print(f"{n} requests / {total_time*1000:.0f} ms = {total_toks/total_time:.0f} tok/s throughput")
asyncio.run(main(32))4. v1 → v2 Performance Comparison
Typical numbers on an A100 40GB:
| Metric | v1 (transformers) | v2 (vLLM) | Gain |
|---|---|---|---|
| Single-request TPOT (ms/tok) | 35–40 | 18–22 | 2× |
| Single-request TTFT (ms) | 800 | 200 | 4× |
| Throughput at 32 concurrency (tok/s) | ≈ 30 (effectively serialized) | 800–1200 | 30–40× |
| Peak memory | 14.5 GB | 36 GB (including KV pool) | Higher, but it buys throughput |
| GPU utilization | 30–50% | 80–95% | Saturated |
Why vLLM is 2× faster even on a single request
vLLM beats transformers even for a single request, and the reason is not PagedAttention (that's a multi-request mechanism) but:
- An optimized attention kernel: vLLM uses FlashAttention2 by default, whose memory traffic is far below the naive implementation;
- CUDA Graphs: the entire decode step is compiled into a graph, eliminating kernel launch overhead;
- A tighter KV cache layout: even for a single request it is faster than transformers' list-of-tensors.
See Kernel Fusion and Custom Kernels and vLLM and PagedAttention.
5. v3: AWQ Quantization + EAGLE Speculative Decoding + Triton
v2 already pushed throughput past 1000 tok/s. The next step is "cut cost + cut latency" — and v3's two weapons are quantization and speculative decoding.
1. AWQ 4-Bit Quantization
AWQ (Activation-aware Weight Quantization) compresses weights to INT4 while activations are still computed in FP16 — the representative weight-only quantization scheme. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.
python
# Step 1: quantize the model to INT4 with autoawq
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model_path = "meta-llama/Llama-2-7b-chat-hf"
quant_path = "Llama-2-7b-chat-hf-awq"
quant_config = { "zero_point": True, "q_group_size": 128, "w_bit": 4, "version": "GEMM" }
model = AutoAWQForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model.quantize(tokenizer, quant_config=quant_config)
model.save_quantized(quant_path)bash
# Step 2: vLLM loads the AWQ-quantized model directly
vllm serve ./Llama-2-7b-chat-hf-awq \
--quantization awq \
--dtype float16 \
--max-model-len 4096 \
--max-num-seqs 256 \
--gpu-memory-utilization 0.9 \
--port 8000Expected changes after AWQ:
| Metric | v2 (bf16) | v3a (AWQ INT4) | Change |
|---|---|---|---|
| Weight memory | 14 GB | 4 GB | 3.5× compression |
| Single-request TPOT | 20 ms | 14 ms | 1.4× speedup (memory-bound benefits) |
| Throughput at 32 concurrency | 1000 tok/s | 1700 tok/s | 1.7× |
| Model quality (MMLU) | 45.3 | 44.8 | −0.5 (essentially lossless) |
What you must do before quantizing
Run an evaluation set to verify accuracy. PTQ quantization schemes like AWQ/GPTQ are a trade-off, not a free lunch — some models (those with many activation outliers) lose accuracy badly after quantization. Minimum validation: run MMLU once, or your business's real evaluation set. See "accuracy not aligned before and after quantization" in Common Pitfalls and Anti-Patterns.
2. EAGLE-3 Speculative Decoding
Speculative decoding uses a small model to guess several tokens, which the large model then verifies in one pass — significantly lowering TPOT at small batch sizes. EAGLE-3 is one of the best-performing schemes of 2024; see Speculative Decoding and Medusa/EAGLE.
bash
# Add the EAGLE-3 draft model when starting vLLM
vllm serve ./Llama-2-7b-chat-hf-awq \
--quantization awq \
--speculative-model "lmsys/Falcon-EAGLE-3-7B" \
--num-speculative-tokens 5 \
--speculative-draft-tensor-parallel-size 1 \
--max-model-len 4096 \
--max-num-seqs 16 \
--port 8000--num-speculative-tokens 5 means the small model guesses 5 tokens at a time and the large model verifies them in one pass — if all guesses are right, a single step emits 6 tokens.
Speculative decoding only pays off at small batch sizes
--max-num-seqs 16 is deliberately small. At large batch sizes (≥ 64), speculative decoding is usually a net loss: the large model is already compute-saturated by the batch, and speculative decoding only adds memory and compute. This is one of the most common anti-patterns — see "speculative decoding at large batch" in Common Pitfalls and Anti-Patterns.
3. Triton Deployment
Production environments rarely have just one model or one version — you need a unified orchestration layer. Triton Inference Server is NVIDIA's standard answer; see Triton Inference Server.
A minimal Triton configuration (vLLM backend + AWQ model):
model_repository/
└── llama2-7b-awq/
├── config.pbtxt
└── 1/
└── model.py # Python backend calling vLLMconfig.pbtxt:
name: "llama2-7b-awq"
backend: "python"
max_batch_size: 32
input [
{ name: "prompt", data_type: TYPE_STRING, dims: [ -1 ] }
]
output [
{ name: "text", data_type: TYPE_STRING, dims: [ -1 ] }
]
dynamic batching {
preferred_batch_size: [ 4, 8, 16, 32 ]
max_queue_delay_microseconds: 50000
}Start Triton:
bash
tritonserver --model-repository ./model_repository \
--http-port 8000 \
--grpc-port 8001 \
--metrics-port 8002Triton's value only fully shows in v3: multiple models coexisting (LLM + embedder + reranker), multi-version gray releases (v2 and v3 running side by side with weighted traffic splitting), and a unified metrics port (/metrics exposing Prometheus metrics). This is the core practice of Model Serving and Orchestration.
6. Performance Comparison Across the Three Versions
Here is the consolidated metrics table (A100 40GB, single GPU, Llama-2-7B-Chat, prompt 200 tokens / output 256 tokens / 32 concurrent requests):
| Metric | v1 PyTorch | v2 vLLM bf16 | v3 vLLM AWQ+EAGLE+Triton | vs. v1 |
|---|---|---|---|---|
| Single-request TTFT (ms) | 800 | 200 | 120 | 6.7× |
| Single-request TPOT (ms/tok) | 38 | 20 | 10 | 3.8× |
| Single-request end-to-end (s) | 10.5 | 5.3 | 2.7 | 3.9× |
| Throughput at 32 concurrency (tok/s) | ≈ 30 | 1100 | 1900 | 63× |
| Peak memory (GB) | 14.5 | 36 | 22 | 14 GB saved |
| GPU utilization | 30–50% | 80–95% | 85–95% | Saturated |
| Model quality (MMLU) | 45.3 | 45.3 | 44.8 | −0.5 |
| Deployment complexity | Low | Medium | High | - |
How to read these performance numbers
- Absolute numbers vary with hardware, prompts, and concurrency: the numbers here are reference values from an A100 40GB. The v1/v2 ratios are similar on an RTX 4090, but the absolute values differ. To reproduce, use the methodology from Inference Benchmarking in Practice.
- v3's cost is complexity: one more quantized model, one more draft model, one more Triton orchestration layer — the on-call debugging chain is longer when things break in production. Whether v3 is worth it depends on your SLA and cost math; see Deployment Design Principles.
- Quality cannot be judged by MMLU alone: business evaluation (accuracy on your real prompt distribution) is the final judge — see Common Pitfalls and Anti-Patterns.
7. Project Directory Structure
A maintainable inference deployment project keeps its files organized around separation of models, code, and configuration:
llama-deploy/
├── README.md # Deployment guide, benchmarks, rollback procedure
├── requirements.txt # Pinned versions
├── configs/
│ ├── v1_transformers.yaml # One config per version
│ ├── v2_vllm.yaml
│ └── v3_vllm_awq_eagle.yaml
├── models/
│ ├── llama2-7b-chat/ # Original weights (read-only)
│ ├── llama2-7b-chat-awq/ # Quantized weights
│ └── falcon-eagle-3-7b/ # Draft model for speculative decoding
├── src/
│ ├── __init__.py
│ ├── client.py # Unified client (OpenAI / Triton)
│ ├── bench.py # Benchmark script
│ └── quantize.py # AWQ quantization script
├── deploy/
│ ├── triton/
│ │ ├── model_repository/
│ │ └── start_triton.sh
│ └── docker/
│ ├── vllm.Dockerfile # vLLM image
│ └── triton.Dockerfile
├── scripts/
│ ├── bench_v1.sh # Run v1 benchmark
│ ├── bench_v2.sh
│ └── bench_v3.sh
└── reports/
└── bench_2024-08-22.md # Three-version performance comparison report| File/Directory | Responsibility | Why it's placed here |
|---|---|---|
configs/ | All three versions' configs in one place | Switching versions only changes launch config, never code |
models/ | Weights decoupled from the training framework | Quantized artifacts stay separate from original weights; rollback is just a config switch |
src/bench.py | Unified benchmarking | All three versions use the same test script, so comparisons are fair |
deploy/ | Deployment concerns isolated | Docker/Triton configs don't pollute the main codebase |
The golden test of a reproducible deployment: delete models/ and reports/, then run scripts/bench_v2.sh on a fresh machine — can you reproduce v2's numbers? If yes, your engineering passes.
8. Common Pitfalls
Pitfall 1: Treating v1's GPU utilization as the only metric
v1 shows 30% GPU utilization, which looks "mostly idle," so a beginner concludes "just add two more concurrent requests." But v1's bottleneck is memory bandwidth, not compute — adding concurrency squeezes the bandwidth harder and throughput actually drops. This is the classic memory-bound scenario from The Roofline Model and Compute Analysis. Get KV cache under control with PagedAttention first; only then talk about adding concurrency.
Pitfall 2: Setting v2's gpu-memory-utilization too high
vLLM defaults to 0.9 (90% of GPU memory for the KV pool). If an embedder or reranker shares the same GPU, 0.9 will OOM the other models. In production, use 0.7–0.8 and leave headroom — see Tuning and Performance Optimization.
Pitfall 3: Skipping the business evaluation for v3's AWQ model
MMLU only measures general capability. If your business is code generation, legal Q&A, or medical summarization, you must validate on a business evaluation set. One team shipped an AWQ code model and watched the JSON output format break outright (INT4 shifted the logits of specific tokens); MMLU barely moved while the business collapsed. See Common Pitfalls and Anti-Patterns.
Pitfall 4: Not counting the EAGLE draft model in the memory budget
EAGLE-3 7B alone takes 14 GB in bf16 — the "small model" is actually as large as the main model. The draft model must be included in the memory budget, or the server OOMs at startup.
9. Next Steps
Finishing the three versions is only the starting point. To grow it into a project you can proudly show, go in four directions:
- Add distributed serving: upgrade single-GPU vLLM to tensor parallelism (TP=2/4) or pipeline parallelism (PP) to serve large models across multiple GPUs — see Distributed Inference (TP/PP).
- Add monitoring: Prometheus + Grafana tracking QPS, p99, KV cache hit rate, and GPU temperature — a deployment without monitoring is running naked.
- Add routing: use a small model for easy requests and a large model for hard ones to cut average cost — see the multi-model routing project in Portfolio Projects.
- Add RAG: wire embedder + vector database + reranker + LLM into an end-to-end RAG service — see Portfolio Projects.
If you want to learn each optimization step at finer granularity, the progressive tutorial from v0 to v5 is in Progressive Tutorial: Three Working Versions. If you want to turn this into a project you can explain clearly in interviews, the expansion routes are in Portfolio Projects and Skills Benchmarking: What to Highlight on Your Resume.
10. Further Reading
- What Is Inference Acceleration? — the theoretical version of the inference optimization landscape
- Anatomy of the Overall Architecture — scaling this article's small project up to a production-grade system
- vLLM and PagedAttention — a deep dive into v2's core engine
- TensorRT-LLM — the extreme-performance alternative to vLLM
- Model Quantization Fundamentals — the principles and comparison of AWQ/GPTQ
- Speculative Decoding and Medusa/EAGLE — the principles behind EAGLE/Medusa
- Progressive Tutorial: Three Working Versions — the fine-grained tutorial version of this article
- Inference Benchmarking in Practice — the methodology behind the three-version comparison table
- Tuning and Performance Optimization — the complete vLLM tuning checklist
- Deployment Design Principles — 12 disciplines before going live
- Common Pitfalls and Anti-Patterns — the full version of Section 8 of this article
References
- vLLM Documentation: Quickstart — official docs for the vLLM launch CLI and its parameters
- vLLM Documentation: Quantization — how AWQ/GPTQ models are loaded in vLLM
- vLLM Documentation: Speculative Decoding — EAGLE/Medusa speculative decoding configuration
- AutoAWQ README — official implementation of AWQ quantization
- Triton Inference Server User Guide — official Triton deployment documentation
- NVIDIA TensorRT-LLM — the extreme-performance path that replaces vLLM
- HuggingFace Transformers: Generate — the
generateinterface used by v1 - NVIDIA Nsight Systems — GPU profiling tool