Skip to content

Deploy an Inference Service from Scratch

At a glance Walk through a three-stage progressive path with Llama-2-7B — from raw PyTorch inference, to swapping in vLLM, to AWQ quantization + EAGLE speculative decoding + Triton deployment. Each stage ships runnable commands, observed metrics, a directory layout, and performance comparisons: the most concentrated end-to-end hands-on project in inference deployment.

Deploy an Inference Service from Scratch ​

Reading ten benchmark reports is worth less than carrying one deployment through end to end. This article walks you through the full journey of an inference service with the classic Llama-2-7B — from "it runs" to "it runs fast" to "it runs steady." This is not about calling a library and getting a number: it explains why each version is built the way it is, where the gains come from, and what the trade-offs cost.

Many beginners, after learning a pile of inference optimization techniques, still don't know what a production-grade inference service actually looks like. How do you lift v1 straight out of the training framework? Does swapping vLLM in for transformers really buy you several times the performance? Will quantization and speculative decoding fight each other when stacked? This article runs the complete deployment flow from start to finish along a single throughline: a Llama-2-7B chat service. What you walk away with is not "this model's tokens/s number" but an optimization skeleton you can reuse for any LLM deployment.

The full project ships in three versions (v1 → v2 → v3). Here is the map first:

        Starting point: Llama-2-7B weights on Hugging Face
                          │
                          ▼
        ┌─────────────────────────────────────────┐
        │ v1: PyTorch + Transformers, direct      │
        │   - transformers from_pregressive       │
        │   - bf16, single batch, no KV cache     │
        │     reuse                               │
        │   - Observe latency / memory baseline   │
        └──────────────────┬──────────────────────┘
                           ▼
        ┌─────────────────────────────────────────┐
        │ v2: vLLM replaces transformers          │
        │   - PagedAttention + continuous batching │
        │   - OpenAI-compatible API service        │
        │   - 5–10× throughput, p99 latency down   │
        └──────────────────┬──────────────────────┘
                           ▼
        ┌─────────────────────────────────────────┐
        │ v3: AWQ INT4 + EAGLE-3 + Triton         │
        │   - 4-bit quantization compresses weights │
        │   - Speculative decoding for extreme      │
        │     latency at small batch                │
        │   - Triton orchestrates multiple models   │
        │     and versions                          │
        └─────────────────────────────────────────┘

The three versions answer three typical engineering questions: v1 answers "can it run at all," v2 answers "is throughput good enough," and v3 answers "can cost and latency be cut further." They are exactly the standard upgrade path for inference optimization in industry — see Learning Paths: Three Routes and Batching and Request Scheduling.

Prerequisites

This article assumes you already understand the basics of LLM inference: Latency, Throughput, and Concurrency, The GPU Memory Hierarchy and the Bandwidth Wall, Model Quantization Fundamentals, and Model Serving and Orchestration. If you still need to catch up, read What Is Inference Acceleration? and Anatomy of the Overall Architecture first, then come back.

1. Choosing the Project: Why Llama-2-7B ​

The first question: which model should you pick for a hands-on project? There are three criteria, and all three are mandatory.

CriterionWhyCost of violating it
Fits on a single GPU (~≤ 14B params)One A100/L40/4090 runs the whole flow without tensor-parallel headachesMulti-GPU distributed inflates the complexity 10× — you lose the main thread
Broad engine support (vLLM/TRT-LLM/llama.cpp all support it)Enables cross-engine comparison — comparison is the shortcut to learningPick an obscure model with no engine support and half the optimization techniques can't be verified
A public chat version availableLets you run a real workload (a Q&A service), not just generate demosWith only a base model, output quality is poor and there is no end-to-end experience

Against these criteria, the three candidates:

ModelParamsSingle-GPU memory (bf16 + KV)Ecosystem supportFit
Llama-2-7B-Chat7B≈ 16 GB★★★ Full coverage: vLLM/TRT-LLM/llama.cpp/Ollama★★★ First choice
Qwen2.5-7B-Instruct7B≈ 16 GB★★★ The most complete ecosystem among Chinese models★★★ First choice for Chinese scenarios
Mistral-7B-Instruct-v0.37B≈ 16 GB★★☆ Supported by vLLM/TRT-LLM★★☆ English scenarios

This article picks Llama-2-7B-Chat for three reasons:

  1. Ecosystem benchmark: vLLM, TensorRT-LLM, and llama.cpp — the three mainstream engines — all treat it as a first-class citizen, so comparisons are fair.
  2. Manageable memory: bf16 inference takes about 14 GB of weights + 2 GB of KV cache, so a single A100 (40 GB) or RTX 4090 (24 GB) can carry all three versions of the experiment.
  3. Publicly downloadable: Meta's official weights are public on Hugging Face (gated access required) — no special permission needed.

Define "success" before writing any code

When you translate a problem into a deployment task, you must answer three questions:

  1. Workload type: online chat (streaming generation, low latency) or offline batch (high throughput)? This article takes online chat as the main thread because it is the most typical and the most demanding.
  2. SLA definition: time to first token (TTFT) < 500 ms, per-request generation latency (TPOT) < 50 ms/token, p99 end-to-end < 4 s, and > 20 QPS on a single GPU.
  3. Baseline: v1's raw PyTorch inference is the baseline. Every later optimization must clearly beat it, or the optimization is pointless.

The SLA is written down before the code. That is basic hygiene from Deployment Design Principles.

2. Setting Up the Environment ​

1. Version Requirements ​

The code in this article is based on the following versions (stable releases from 2024–2025):

SoftwareVersionPurpose
Python3.10+The language itself
torch2.3+ / CUDA 12.1Training framework and GPU backend
transformers4.44+Hugging Face model loading
vllm0.6+Primary inference engine for v2/v3
autoawq0.2.6+AWQ 4-bit quantization
tritonserver24.08+Triton inference server
nvidia-smidriver ≥ 535GPU monitoring

2. Create a Virtual Environment ​

Never install dependencies into the system Python. Inference optimization is extremely version-sensitive — a one-digit mismatch in CUDA/cuDNN/torch versions can cost you 30% of performance:

bash
mkdir llama-deploy && cd llama-deploy

python -m venv .venv
source .venv/bin/activate        # macOS / Linux
# .venv\Scripts\activate        # Windows PowerShell

# v1 uses: transformers + torch
pip install torch==2.3.1 transformers==4.44.2 accelerate sentencepiece

# v2 uses: vLLM (bundles PagedAttention)
pip install vllm==0.6.3

# v3 uses: AWQ quantization + Triton client
pip install autoawq==0.2.6 tritonclient[all]==2.44.0

pip freeze > requirements.txt

Why vLLM gets its own install

vLLM's wheels ship precompiled CUDA kernels (PagedAttention, FlashAttention) that demand an exact match between CUDA and torch versions. Mixed into the transformers environment, they conflict constantly. In production practice, v1 and v2 are usually split into two separate virtual environments or two separate Docker images — which is exactly why Triton later became the "unified orchestration layer." See Triton Inference Server.

3. GPU and Memory Self-Check ​

bash
nvidia-smi --query-gpu=name,memory.total,memory.free,driver_version,compute_cap --format=csv
# name, memory.total [MiB], memory.free [MiB], driver_version, compute_cap
# NVIDIA A100 80GB, 81920, 78000, 535.104.05, 8.0

compute_cap must be ≥ 8.0 (Ampere or newer) to run this article's full bf16 + PagedAttention + FlashAttention path end to end. Older cards like the T4/V100 can still run it, but their performance numbers are in a different class.

3. v1: Raw PyTorch + Transformers Inference ​

v1 has exactly one goal: make the model produce decent answers. It deliberately uses no tricks; it is the control baseline for everything that follows.

1. Load the Model ​

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "meta-llama/Llama-2-7b-chat-hf"

tokenizer = AutoTokenizer.from_pretrained(MODEL, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
    MODEL,
    torch_dtype=torch.bfloat16,   # bf16 is more numerically stable than fp16 and halves memory vs fp32
    device_map="auto",            # automatically move weights onto the GPU
)
model.eval()

prompt = "Explain PagedAttention in three sentences."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

device_map="auto" is a capability of Hugging Face accelerate — it distributes model layers across the visible GPUs/CPU according to the memory budget. On a single GPU, it simply means "move everything to the GPU."

2. Run One Inference ​

python
with torch.inference_mode():
    out = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=True,
        temperature=0.6,
        top_p=0.9,
        pad_token_id=tokenizer.eos_token_id,
    )
text = tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(text)

3. Observe the Baseline Metrics ​

The point of v1 is not "it runs" — it is measuring the baseline: single-request latency, peak memory, and tokens/s. This is the reference point for every optimization that follows.

python
import time, torch

def bench_one(prompt, max_new_tokens=256):
    torch.cuda.reset_peak_memory_stats()
    inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
    # warmup (the first inference triggers kernel compilation and must not enter the baseline)
    _ = model.generate(**inputs, max_new_tokens=8, do_sample=False)

    torch.cuda.synchronize()
    t0 = time.perf_counter()
    out = model.generate(**inputs, max_new_tokens=max_new_tokens,
                         do_sample=True, temperature=0.6, top_p=0.9,
                         pad_token_id=tokenizer.eos_token_id)
    torch.cuda.synchronize()
    t1 = time.perf_counter()

    peak = torch.cuda.max_memory_allocated() / 1e9   # GB
    n_tok = out.shape[1] - inputs["input_ids"].shape[1]
    tps = n_tok / (t1 - t0)
    print(f"Output {n_tok} tokens / {(t1-t0)*1000:.0f} ms = {tps:.1f} tok/s | peak {peak:.1f} GB")
    return {"tokens": n_tok, "ms": (t1-t0)*1000, "tps": tps, "peak_gb": peak}

bench_one("Write a 200-word essay about autumn.")

Typical v1 baseline on an A100:

MetricValueNote
Tokens generated per request256Aligned with max_new_tokens
End-to-end latency≈ 9–11 sIncludes prefill + decode
Decode-phase tokens/s≈ 25–30 tok/sSingle batch, no PagedAttention
Peak memory≈ 14.5 GB7B weights 14 GB + KV cache + temporaries
GPU utilization30–50%Decode phase is severely memory-bound

That number is the baseline. Write it down — v2 and v3 will both be measured against it. Why are latency and GPU utilization so low? See The Roofline Model and Compute Analysis and The GPU Memory Hierarchy and the Bandwidth Wall.

The two core problems v1 exposes

  1. Extremely low throughput: 30 tok/s means a single GPU can serve 30 concurrent users at 1 token/s each — nowhere near a SaaS bar.
  2. Low GPU utilization: 30–50% means compute is sitting idle. The reason: Hugging Face's generate decodes token by token by default with batch=1 and a primitive KV cache reuse mechanism, so every step is limited by memory bandwidth, not compute — which is exactly the core problem Batching and Request Scheduling exists to solve.

4. v2: vLLM Replaces Transformers ​

v2's goal is to pull throughput up with vLLM. Its two core mechanisms — PagedAttention and continuous batching — are the biggest engineering breakthroughs in inference optimization over the past two years. See vLLM and PagedAttention.

1. Start an OpenAI-Compatible Server with the CLI ​

vLLM ships with the vllm serve CLI, which brings up an OpenAI-protocol HTTP server in a few lines:

bash
# Start Llama-2-7B-Chat on a single GPU with the OpenAI-compatible API
vllm serve meta-llama/Llama-2-7b-chat-hf \
    --dtype bfloat16 \
    --max-model-len 4096 \
    --max-num-seqs 256 \
    --gpu-memory-utilization 0.9 \
    --enforce-eager \
    --port 8000

What the key parameters mean (see Tuning and Performance Optimization):

ParameterMeaningv1 equivalent
--max-model-lenMaximum context length of the modelThe max_new_tokens cap in transformers
--max-num-seqsMaximum number of requests in flightNone in v1 — effectively batch=1
--gpu-memory-utilizationFraction of total GPU memory used for the KV cacheNo active management in v1
--enforce-eagerDisable CUDA Graphs (for debugging)-

2. Client Test ​

python
from openai import OpenAI
import time, asyncio

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

def bench_one_vllm(prompt, max_tokens=256):
    t0 = time.perf_counter()
    resp = client.completions.create(
        model="meta-llama/Llama-2-7b-chat-hf",
        prompt=prompt,
        max_tokens=max_tokens,
        temperature=0.6,
        top_p=0.9,
    )
    t1 = time.perf_counter()
    n_tok = resp.usage.completion_tokens
    print(f"Output {n_tok} tokens / {(t1-t0)*1000:.0f} ms = {n_tok/(t1-t0):.1f} tok/s")
    return {"tokens": n_tok, "ms": (t1-t0)*1000, "tps": n_tok/(t1-t0)}

bench_one_vllm("Write a 200-word essay about autumn.")

3. Concurrent Load Test ​

v1 was slow even for a single request; vLLM's real value shows under concurrent requests. Below, 32 requests hit the server at once:

python
import asyncio
from openai import AsyncOpenAI

aclient = AsyncOpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

async def one(i):
    t0 = time.perf_counter()
    r = await aclient.completions.create(
        model="meta-llama/Llama-2-7b-chat-hf",
        prompt=f"Explain PagedAttention in three sentences. Request number {i}.",
        max_tokens=128,
    )
    return r.usage.completion_tokens, time.perf_counter() - t0

async def main(n=32):
    t0 = time.perf_counter()
    results = await asyncio.gather(*[one(i) for i in range(n)])
    total_toks = sum(r[0] for r in results)
    total_time = time.perf_counter() - t0
    print(f"{n} requests / {total_time*1000:.0f} ms = {total_toks/total_time:.0f} tok/s throughput")

asyncio.run(main(32))

4. v1 → v2 Performance Comparison ​

Typical numbers on an A100 40GB:

Metricv1 (transformers)v2 (vLLM)Gain
Single-request TPOT (ms/tok)35–4018–222×
Single-request TTFT (ms)8002004×
Throughput at 32 concurrency (tok/s)≈ 30 (effectively serialized)800–120030–40×
Peak memory14.5 GB36 GB (including KV pool)Higher, but it buys throughput
GPU utilization30–50%80–95%Saturated

Why vLLM is 2× faster even on a single request

vLLM beats transformers even for a single request, and the reason is not PagedAttention (that's a multi-request mechanism) but:

  1. An optimized attention kernel: vLLM uses FlashAttention2 by default, whose memory traffic is far below the naive implementation;
  2. CUDA Graphs: the entire decode step is compiled into a graph, eliminating kernel launch overhead;
  3. A tighter KV cache layout: even for a single request it is faster than transformers' list-of-tensors.

See Kernel Fusion and Custom Kernels and vLLM and PagedAttention.

5. v3: AWQ Quantization + EAGLE Speculative Decoding + Triton ​

v2 already pushed throughput past 1000 tok/s. The next step is "cut cost + cut latency" — and v3's two weapons are quantization and speculative decoding.

1. AWQ 4-Bit Quantization ​

AWQ (Activation-aware Weight Quantization) compresses weights to INT4 while activations are still computed in FP16 — the representative weight-only quantization scheme. See Model Quantization Fundamentals and Weight-Only Quantization and Mixed Precision.

python
# Step 1: quantize the model to INT4 with autoawq
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer

model_path = "meta-llama/Llama-2-7b-chat-hf"
quant_path = "Llama-2-7b-chat-hf-awq"
quant_config = { "zero_point": True, "q_group_size": 128, "w_bit": 4, "version": "GEMM" }

model = AutoAWQForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model.quantize(tokenizer, quant_config=quant_config)
model.save_quantized(quant_path)
bash
# Step 2: vLLM loads the AWQ-quantized model directly
vllm serve ./Llama-2-7b-chat-hf-awq \
    --quantization awq \
    --dtype float16 \
    --max-model-len 4096 \
    --max-num-seqs 256 \
    --gpu-memory-utilization 0.9 \
    --port 8000

Expected changes after AWQ:

Metricv2 (bf16)v3a (AWQ INT4)Change
Weight memory14 GB4 GB3.5× compression
Single-request TPOT20 ms14 ms1.4× speedup (memory-bound benefits)
Throughput at 32 concurrency1000 tok/s1700 tok/s1.7×
Model quality (MMLU)45.344.8−0.5 (essentially lossless)

What you must do before quantizing

Run an evaluation set to verify accuracy. PTQ quantization schemes like AWQ/GPTQ are a trade-off, not a free lunch — some models (those with many activation outliers) lose accuracy badly after quantization. Minimum validation: run MMLU once, or your business's real evaluation set. See "accuracy not aligned before and after quantization" in Common Pitfalls and Anti-Patterns.

2. EAGLE-3 Speculative Decoding ​

Speculative decoding uses a small model to guess several tokens, which the large model then verifies in one pass — significantly lowering TPOT at small batch sizes. EAGLE-3 is one of the best-performing schemes of 2024; see Speculative Decoding and Medusa/EAGLE.

bash
# Add the EAGLE-3 draft model when starting vLLM
vllm serve ./Llama-2-7b-chat-hf-awq \
    --quantization awq \
    --speculative-model "lmsys/Falcon-EAGLE-3-7B" \
    --num-speculative-tokens 5 \
    --speculative-draft-tensor-parallel-size 1 \
    --max-model-len 4096 \
    --max-num-seqs 16 \
    --port 8000

--num-speculative-tokens 5 means the small model guesses 5 tokens at a time and the large model verifies them in one pass — if all guesses are right, a single step emits 6 tokens.

Speculative decoding only pays off at small batch sizes

--max-num-seqs 16 is deliberately small. At large batch sizes (≥ 64), speculative decoding is usually a net loss: the large model is already compute-saturated by the batch, and speculative decoding only adds memory and compute. This is one of the most common anti-patterns — see "speculative decoding at large batch" in Common Pitfalls and Anti-Patterns.

3. Triton Deployment ​

Production environments rarely have just one model or one version — you need a unified orchestration layer. Triton Inference Server is NVIDIA's standard answer; see Triton Inference Server.

A minimal Triton configuration (vLLM backend + AWQ model):

model_repository/
└── llama2-7b-awq/
    ├── config.pbtxt
    └── 1/
        └── model.py       # Python backend calling vLLM

config.pbtxt:

name: "llama2-7b-awq"
backend: "python"
max_batch_size: 32
input [
  { name: "prompt", data_type: TYPE_STRING, dims: [ -1 ] }
]
output [
  { name: "text", data_type: TYPE_STRING, dims: [ -1 ] }
]
dynamic batching {
  preferred_batch_size: [ 4, 8, 16, 32 ]
  max_queue_delay_microseconds: 50000
}

Start Triton:

bash
tritonserver --model-repository ./model_repository \
             --http-port 8000 \
             --grpc-port 8001 \
             --metrics-port 8002

Triton's value only fully shows in v3: multiple models coexisting (LLM + embedder + reranker), multi-version gray releases (v2 and v3 running side by side with weighted traffic splitting), and a unified metrics port (/metrics exposing Prometheus metrics). This is the core practice of Model Serving and Orchestration.

6. Performance Comparison Across the Three Versions ​

Here is the consolidated metrics table (A100 40GB, single GPU, Llama-2-7B-Chat, prompt 200 tokens / output 256 tokens / 32 concurrent requests):

Metricv1 PyTorchv2 vLLM bf16v3 vLLM AWQ+EAGLE+Tritonvs. v1
Single-request TTFT (ms)8002001206.7×
Single-request TPOT (ms/tok)3820103.8×
Single-request end-to-end (s)10.55.32.73.9×
Throughput at 32 concurrency (tok/s)≈ 301100190063×
Peak memory (GB)14.5362214 GB saved
GPU utilization30–50%80–95%85–95%Saturated
Model quality (MMLU)45.345.344.8−0.5
Deployment complexityLowMediumHigh-

How to read these performance numbers

  1. Absolute numbers vary with hardware, prompts, and concurrency: the numbers here are reference values from an A100 40GB. The v1/v2 ratios are similar on an RTX 4090, but the absolute values differ. To reproduce, use the methodology from Inference Benchmarking in Practice.
  2. v3's cost is complexity: one more quantized model, one more draft model, one more Triton orchestration layer — the on-call debugging chain is longer when things break in production. Whether v3 is worth it depends on your SLA and cost math; see Deployment Design Principles.
  3. Quality cannot be judged by MMLU alone: business evaluation (accuracy on your real prompt distribution) is the final judge — see Common Pitfalls and Anti-Patterns.

7. Project Directory Structure ​

A maintainable inference deployment project keeps its files organized around separation of models, code, and configuration:

llama-deploy/
├── README.md                    # Deployment guide, benchmarks, rollback procedure
├── requirements.txt             # Pinned versions
├── configs/
│   ├── v1_transformers.yaml     # One config per version
│   ├── v2_vllm.yaml
│   └── v3_vllm_awq_eagle.yaml
├── models/
│   ├── llama2-7b-chat/          # Original weights (read-only)
│   ├── llama2-7b-chat-awq/      # Quantized weights
│   └── falcon-eagle-3-7b/       # Draft model for speculative decoding
├── src/
│   ├── __init__.py
│   ├── client.py                # Unified client (OpenAI / Triton)
│   ├── bench.py                 # Benchmark script
│   └── quantize.py              # AWQ quantization script
├── deploy/
│   ├── triton/
│   │   ├── model_repository/
│   │   └── start_triton.sh
│   └── docker/
│       ├── vllm.Dockerfile      # vLLM image
│       └── triton.Dockerfile
├── scripts/
│   ├── bench_v1.sh              # Run v1 benchmark
│   ├── bench_v2.sh
│   └── bench_v3.sh
└── reports/
    └── bench_2024-08-22.md      # Three-version performance comparison report
File/DirectoryResponsibilityWhy it's placed here
configs/All three versions' configs in one placeSwitching versions only changes launch config, never code
models/Weights decoupled from the training frameworkQuantized artifacts stay separate from original weights; rollback is just a config switch
src/bench.pyUnified benchmarkingAll three versions use the same test script, so comparisons are fair
deploy/Deployment concerns isolatedDocker/Triton configs don't pollute the main codebase

The golden test of a reproducible deployment: delete models/ and reports/, then run scripts/bench_v2.sh on a fresh machine — can you reproduce v2's numbers? If yes, your engineering passes.

8. Common Pitfalls ​

Pitfall 1: Treating v1's GPU utilization as the only metric ​

v1 shows 30% GPU utilization, which looks "mostly idle," so a beginner concludes "just add two more concurrent requests." But v1's bottleneck is memory bandwidth, not compute — adding concurrency squeezes the bandwidth harder and throughput actually drops. This is the classic memory-bound scenario from The Roofline Model and Compute Analysis. Get KV cache under control with PagedAttention first; only then talk about adding concurrency.

Pitfall 2: Setting v2's gpu-memory-utilization too high ​

vLLM defaults to 0.9 (90% of GPU memory for the KV pool). If an embedder or reranker shares the same GPU, 0.9 will OOM the other models. In production, use 0.7–0.8 and leave headroom — see Tuning and Performance Optimization.

Pitfall 3: Skipping the business evaluation for v3's AWQ model ​

MMLU only measures general capability. If your business is code generation, legal Q&A, or medical summarization, you must validate on a business evaluation set. One team shipped an AWQ code model and watched the JSON output format break outright (INT4 shifted the logits of specific tokens); MMLU barely moved while the business collapsed. See Common Pitfalls and Anti-Patterns.

Pitfall 4: Not counting the EAGLE draft model in the memory budget ​

EAGLE-3 7B alone takes 14 GB in bf16 — the "small model" is actually as large as the main model. The draft model must be included in the memory budget, or the server OOMs at startup.

9. Next Steps ​

Finishing the three versions is only the starting point. To grow it into a project you can proudly show, go in four directions:

  1. Add distributed serving: upgrade single-GPU vLLM to tensor parallelism (TP=2/4) or pipeline parallelism (PP) to serve large models across multiple GPUs — see Distributed Inference (TP/PP).
  2. Add monitoring: Prometheus + Grafana tracking QPS, p99, KV cache hit rate, and GPU temperature — a deployment without monitoring is running naked.
  3. Add routing: use a small model for easy requests and a large model for hard ones to cut average cost — see the multi-model routing project in Portfolio Projects.
  4. Add RAG: wire embedder + vector database + reranker + LLM into an end-to-end RAG service — see Portfolio Projects.

If you want to learn each optimization step at finer granularity, the progressive tutorial from v0 to v5 is in Progressive Tutorial: Three Working Versions. If you want to turn this into a project you can explain clearly in interviews, the expansion routes are in Portfolio Projects and Skills Benchmarking: What to Highlight on Your Resume.

10. Further Reading ​

References ​