Theme
Deployment and Serving
The goal of deployment is not "getting the model to run," but rather delivering the model to users at acceptable cost, with predictable latency and stable throughput — it's an engineering trade-off around memory, compute, concurrency, and cost.
Getting a model to run on a laptop is just the beginning. Production environments answer completely different questions: Is there enough memory? What's the latency at 100 concurrent requests? Which GPU is most cost-effective? Will quantization drop quality? This article provides the complete deployment chain: fix metrics first, then calculate memory, then select quantization and batching strategies, and finally land on a runnable vLLM example. Underlying mechanisms (autoregressive loop, KV cache, sampling) are at Inference Fundamentals.
I. Fix Metrics First: Deployment Language Is Numbers
The deployment team must speak the same metric language:
| Metric | Full Name / Meaning | Measures What | Typical Scale (Reference) |
|---|---|---|---|
| TTFT | Time To First Token, first-token latency | How long the user "waits for the first response" | Hundreds of ms to seconds |
| TPOT | Time Per Output Token, time per generated token | Generation speed | Tens to hundreds of ms/token |
| Throughput | tokens/s, output volume per unit time | Overall service capacity | Hundreds to thousands of tokens/s per GPU |
| Concurrency | Number of requests served simultaneously | Capacity | Strongly related to memory / batching |
text
User experience conversion: TTFT < 1s is "acceptable"; the faster TPOT, the stronger the "typewriter feel."
Throughput × tokens per request ÷ concurrency ≈ number of GPUs you need (rough estimate).Latency vs. throughput trade-off
Latency-prioritized (interactive dialogue) and throughput-prioritized (offline batch processing) are two completely different optimization directions: the former reduces batch size to preserve speed; the latter increases batch size to maximize throughput. Clarify which one your business belongs to before tuning parameters.
Reference Ranges for Metric Targets and How to Set Them
Metrics don't have absolute "standard values," but reference ranges can be given by business type (actual values depend on testing and product requirements):
| Business Type | TTFT Target | TPOT Target | Primary Constraint |
|---|---|---|---|
| Real-time dialogue (Chat) | < 1s | As fast as possible (typewriter feel) | Latency-first |
| Customer service / ticket assistance | < 2s | < 100ms/token | Latency + quality |
| Code completion | < 500ms | < 50ms/token | Extreme latency sensitivity |
| Offline batch (summarization / extraction) | Not sensitive | Not sensitive | Throughput + cost-first |
Principle for setting targets: first measure the current baseline (P50 and P95), then set an "acceptable P95" rather than targeting the "average" — P95 represents the real user experience; averages are pulled down by a few fast requests.
II. Memory Estimation: Weights + KV Cache + Activations
Inference memory consists of three parts — learn to calculate by hand first, then trust any monitoring tool:
Total Memory ≈ Weight Memory + KV Cache Memory + Activation / Temporary Memory1. Weight Memory
Weight Memory = Parameter Count × Bytes per Parameter
fp16/bf16: 2 bytes/parameter INT8: 1 byte/parameter INT4: 0.5 bytes/parameter| Model | Parameter Count | FP16 | INT8 | INT4 |
|---|---|---|---|---|
| 7B/8B | ~8B | ~16 GB | ~8 GB | ~4 GB |
| 13B | ~13B | ~26 GB | ~13 GB | ~6.5 GB |
| 70B | ~70B | ~140 GB | ~70 GB | ~35 GB |
2. KV Cache Memory
KV cache stores the Key/Value generated at each layer for each sequence (mechanism at Transformer Architecture Explained), and grows linearly with: concurrent requests × context length:
KV memory per token = 2 × num_layers × num_kv_heads × head_dim × bytes_per_elem
KV cache total = KV memory per token × context_length × concurrencyUsing Llama-2 70B (80 layers, 8 KV heads, head dim 128, fp16) as an example:
Per token = 2 × 80 × 8 × 128 × 2 bytes = 327,680 bytes ≈ 0.31 MB/token
4096 context × 32 concurrency ≈ 0.31MB × 4096 × 32 ≈ 40 GB !These numbers reveal the first law of deployment: KV cache grows much faster than weights; with long context + high concurrency, it becomes the dominant memory consumer (this is an area that neither MoE sparse expert models nor quantization can fully address). Engineering optimizations for KV cache are in Section IV.
3. Activation Memory
Intermediate activations at each layer during the prefill phase also consume memory, proportional to batch size and sequence length, typically significantly compressed by fused operators like FlashAttention. For estimation, multiply the "weights + KV cache" result by 1.1–1.3 to leave activation headroom.
The difference between inference and training memory
Training memory also includes gradients and optimizer states (full-parameter training ≈ 12–16 bytes per parameter), 4–8× that of inference. "70B model, 140 GB memory" is the inference metric; training is calculated separately. Training-side estimation is in the QLoRA table at Fine-Tuning in Practice.
III. Quantization: Trading Precision for Memory and Speed
Quantization compresses weights from fp16 to low bit-width — the primary method for "running large models on small cards" and "reducing costs."
| Method | Precision | Characteristics | Typical Tools |
|---|---|---|---|
| Dynamic quantization (INT8) | 8-bit | General-purpose, simple, minimal quality drop | vLLM built-in, bitsandbytes |
| GPTQ | 4-bit | Post-training quantization, dequantizes at activation time; significant memory savings | auto-gptq, supported by vLLM |
| AWQ | 4-bit | Activation-aware quantization, smaller precision loss | Supported by vLLM, autoawq |
| GGUF (K-quant) | 1–8-bit | llama.cpp ecosystem, runs on CPU / edge | llama.cpp, Ollama |
| FP8 | 8-bit | Native support on newer GPUs (H series) | TensorRT-LLM, vLLM |
bash
# Load a quantized model directly with vLLM (supports both AWQ and GPTQ)
pip install vllm
# Load 4-bit AWQ version of Qwen2.5-7B
python -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen2.5-7B-Instruct-AWQ \
--quantization awq \
--gpu-memory-utilization 0.9 \
--max-model-len 8192Quantization is not free
4-bit quantization typically causes a 1–3% drop in benchmark scores (varies by task and model), but the inference speed and memory gains are enormous. Decision sequence: run fp16 baseline first, then run INT8/AWQ comparison, confirm the quality drop is acceptable using Evaluation in Practice methods, then launch.
IV. Batching and KV Cache Optimization
1. Continuous Batching: The Biggest Throughput Lever
Traditional batching is "static" (the next batch starts only after all current requests complete), causing short requests to wait for long ones. Continuous batching lets sequences enter and exit the batch at the token level — as soon as one completes, space is freed for a new request, significantly improving GPU utilization.
2. PagedAttention: KV Cache Memory Management
vLLM's core innovation, PagedAttention, manages KV cache in "pages" (analogous to virtual memory in operating systems), eliminating fragmentation and allocating on demand, increasing throughput 2–4× under the same memory (official benchmarks; actual numbers vary by workload, see References).
3. Quick Overview of Other Optimizations
| Optimization | Effect |
|---|---|
| FlashAttention | Fused operator, saves memory + speeds up |
| Prefix caching | Reuse KV cache for identical system prompts |
| Speculative decoding | Small model drafts + large model verifies, reduces TPOT |
| Structured output constrained decoding | Skip invalid tokens when generating JSON, saves time and stabilizes format |
4. Multi-GPU Deployment: Tensor Parallelism
When a model doesn't fit on a single card, use Tensor Parallelism (TP) to split a layer across multiple GPUs, with GPUs synchronizing via high-speed interconnects (NVLink / InfiniBand). Key points:
| Parallelism Type | What Gets Split | When to Use |
|---|---|---|
| Tensor parallelism | Matrices / attention heads within a layer split across GPUs | When the model doesn't fit on one card; sensitive to inter-GPU bandwidth |
| Pipeline parallelism | Layers grouped to different GPUs | Many layers; slight throughput decrease |
| Data parallelism | Each GPU gets one model copy, data is split | Throughput scaling (used with TP) |
vLLM enables TP with one line: --tensor-parallel-size 2; note that TP requires high-speed inter-GPU connectivity, otherwise communication will tank performance. 70B fp16 commonly uses 2×80GB or 4×40GB configurations.
V. GPU Selection
| GPU | Memory | Reference Positioning (based on 2024–2025 mainstream market; prices fluctuate) |
|---|---|---|
| RTX 4090 | 24 GB | Consumer grade: 7B at full precision / 13B INT4 on one card |
| L40S | 48 GB | 13B fp16 / 70B INT4 on one card |
| A100 40/80 GB | 40/80 GB | Datacenter workhorse: 70B INT4 (80 GB) |
| H100 80 GB | 80 GB | Flagship for large model training / inference, native FP8 |
| H20 / domestic GPUs | Varies | Choice under constrained environments; check ecosystem compatibility |
Selection decision table (how to pick a GPU):
| Your Model + Precision | Minimum Single-Card Requirement | Notes |
|---|---|---|
| 7B INT4 | 8–16 GB is enough | Edge / low-budget |
| 7B fp16 | 24 GB | RTX 4090 class |
| 13B fp16 / 70B INT4 | 40–48 GB | L40S / A100 |
| 70B fp16 | 80 GB × 2 (tensor parallelism) | H100 / A100 |
| High concurrency | Increase memory per "KV cache formula" | Don't just look at weights |
Quick criterion for "is memory enough?"
Weight memory + KV cache (at peak concurrency × peak context) ≤ single-card memory × gpu-memory-utilization. If it exceeds, quantize first; if still not enough, use multi-GPU tensor parallelism or reduce concurrency. Always account for KV cache — accounting only for weights is the #1 source of deployment accidents.
VI. vLLM Deployment Example: End-to-End
vLLM provides an OpenAI-compatible API, so frontend code doesn't need a single line of change.
bash
# ① Start the service (using a 7B model as example)
python -m vllm.entrypoints.openai.api_server \
--model Qwen/Qwen2.5-7B-Instruct \
--served-model-name my-llm \
--gpu-memory-utilization 0.9 \
--max-model-len 8192 \
--enforce-eager # Can disable CUDA graph for environments with poor compatibility
# ② Call via OpenAI SDK (client-side is transparent)python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="my-llm",
messages=[{"role": "user", "content": "What is a KV cache?"}],
temperature=0.7,
max_tokens=512,
stream=True, # Streaming output
)
for chunk in resp:
print(chunk.choices[0].delta.content or "", end="")bash
# ③ Monitoring and load testing
curl http://localhost:8000/v1/models # View model list
# Load testing: use openai-benchmark / h2o-llmstudio or custom concurrency scripts
# Key monitored metrics: TTFT percentiles, TPOT percentiles, GPU utilization, memory usage, queue backlogOpenAI-Compatible API: Protocol Details
vLLM's /v1/chat/completions is compatible with the OpenAI protocol. Four details that commonly trip people up:
| Detail | Description |
|---|---|
| Streaming (stream) | stream=true returns data: fragments (SSE), ending with data: [DONE] |
| Tool calling | tools parameter and tool_calls response fields require client-side support |
| Structured output | JSON mode / response_format support varies by engine and model; test first |
| Auth and exposure | Self-hosted services need their own auth (API keys); don't expose publicly without protection |
python
# SSE handling for streaming responses (client-side)
import requests
resp = requests.post(
"http://localhost:8000/v1/chat/completions",
json={"model": "my-llm",
"messages": [{"role": "user", "content": "Hello"}],
"stream": True},
headers={"Authorization": "Bearer EMPTY"},
stream=True,
)
for line in resp.iter_lines():
if line and line.startswith(b"data: "):
payload = line[6:]
if payload == b"[DONE]":
break
# Parse JSON, take choices[0].delta.content and print fragment by fragmentCompatible ≠ functionally equivalent
"OpenAI compatible" is protocol-compatible, not functionally equivalent: different engines support extended fields like tools, response_format, and logprobs to different degrees. Test every field you plan to use before launch rather than assuming they're available.
Deployment Checklist (Verify Each Item Before Launch)
| Item | Check |
|---|---|
| Memory budget | Weights + KV cache (peak concurrency × peak context) ≤ available memory |
| Quantization validation | INT4 vs. fp16 baseline difference has been evaluated |
| Metrics met | TTFT/TPOT satisfy product requirements (including P95) |
| Compatibility | OpenAI-compatible API has been integrated (streaming / tool calling / JSON mode) |
| Monitoring | Latency, throughput, memory, error rate, queue backlog all connected to alerts |
| Rollback | Model/config versions are versioned and can be rolled back in one click |
VII. Canary Deployment and Monitoring
Launch is not the end — it's the start of a new round of engineering:
| Method | Practice |
|---|---|
| Canary | New model serves only 5% of traffic; observe metrics and user feedback before scaling |
| Shadow mode | Old and new models run in parallel; new model's results are logged but not shown, compared offline |
| Monitoring | TTFT/TPOT P50/P95, error rate, queue length, GPU utilization |
| Evaluation linkage | Online spot-check results flow back into the golden set, forming a closed loop (see Evaluation in Practice) |
Common deployment accident areas
Memory OOM (only counted weights, not KV cache), cascading failures from unlimited retries (set timeouts and backoff), streaming interruption without reconnection, difficult model version rollback. Write "how to fail gracefully" into the design, don't wait for it to happen.
Monitoring Metrics System: What to Watch After Launch
| Category | Metric | Alert Threshold (Example) |
|---|---|---|
| Latency | P50, P95 of TTFT / TPOT | P95 TTFT exceeds 2s |
| Capacity | Concurrent requests, queue backlog | Queue consistently backing up |
| Compute | GPU utilization, memory usage | Memory usage > 95% |
| Quality | Error rate, retry rate, timeout rate | Error rate > 1% |
| Business | Completion rate, negative feedback rate | Negative feedback rate increasing |
The value of monitoring data lies in trends: memory usage slowly creeping from 60% to 90% signals traffic growth or a memory leak — something worth addressing before "the system crashes today."
VIII. Cloud vs. Local vs. API: Three Options for Decision
| Dimension | Self-hosted GPUs (cloud/local) | Managed API (OpenAI/Anthropic, etc.) | Self-hosted open-source |
|---|---|---|---|
| Cost curve | High fixed cost, low marginal cost | Pay-per-use, no sunk cost | Same, but requires ops |
| Data security | Controlled data egress | Data leaves network (evaluate compliance) | Data stays entirely internal |
| Capability ceiling | Limited by open-source models | Closed-source flagships are strongest | Limited by open-source models |
| Ops burden | High (K8s, GPU, elasticity) | Zero | Medium–high |
| Customization | Fully controllable | Limited (fine-tuning needs platform support) | Fully controllable |
Recommended decision sequence:
Data can leave network, need strongest capability, don't want ops → Managed API
Data-sensitive / compliance restrictions / deep customization → Self-host open-source (vLLM)
Large and stable volume → Self-hosted GPUs (or hybrid: API fallback + self-hosted primary)Cost Accounting: A Simplified Example
Do the math before selection (numbers are for teaching only; prices change frequently; refer to official pricing):
Scenario: 7B model, 100K requests/day, avg input 1000 tokens, output 300 tokens
Daily token count ≈ 100K × (1000 + 300) ≈ 130M tokens
Option A: Managed API (assuming $0.004/M input tokens, $0.016/M output tokens)
Cost ≈ 100K × (1000 × $0.004 + 300 × $0.016) / 1M ≈ $0.88/day ≈ $320/year
Option B: Self-host 1×RTX 4090 (~24 GB, can run 7B INT4)
One-time hardware ~$400; electricity/ops extra; capacity ceiling ~tens to hundreds of thousands of tokens/minConclusion pattern: low volume → API (zero ops), high and stable volume → self-host (amortize fixed cost), middle ground → hybrid. Don't forget to include dev/ops labor cost — that's often the most expensive part.
Further Reading
- Inference Fundamentals: Autoregression and Sampling — Underlying principles of KV cache, prefill/decode phases
- Transformer Architecture Explained — Where KV cache comes from: attention mechanism review
- Context Windows and Long Contexts — How long context affects KV cache and memory
- Frameworks and Tooling: A Selection Guide — Selection tables for vLLM/SGLang/TensorRT-LLM
- MoE Sparse Expert Models — Memory and expert caching challenges for MoE models
- Evaluation in Practice — Regression validation for quantization / new version launches
- Common Pitfalls and Anti-Patterns — Deployment-related pitfalls: ignoring costs, context stuffing
References
- vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention (arXiv:2309.06180) — PagedAttention and vLLM system paper
- vLLM Official Documentation — Official manual for deployment, quantization, LoRA, OpenAI-compatible API
- vLLM Blog: Continuous Batching — Official documentation on continuous batching and throughput benchmarks
- Anyscale: Continuous Batching for LLM Inference — Principles and data on continuous batching
- GPTQ: Accurate Post-Training Quantization (arXiv:2210.17323) — Original GPTQ paper
- AWQ: Activation-aware Weight Quantization (arXiv:2306.00978) — Original AWQ paper
- llama.cpp (GitHub) — Official repo for GGUF quantization and CPU inference
- Ollama — One-click local GGUF model runner
- OpenAI API Reference — Protocol reference for OpenAI-compatible APIs
- FlashAttention (arXiv:2205.14135) — Original paper on fused attention operator