Skip to content

Deployment and Serving

At a glance From memory estimation formulas (weights + KV cache + activations), quantization (GPTQ/AWQ/GGUF), continuous batching, to TTFT/TPOT metrics, GPU selection, and a vLLM deployment example — covering the complete pipeline and decision framework for "taking a model from your laptop to production."

Deployment and Serving ​

The goal of deployment is not "getting the model to run," but rather delivering the model to users at acceptable cost, with predictable latency and stable throughput — it's an engineering trade-off around memory, compute, concurrency, and cost.

Getting a model to run on a laptop is just the beginning. Production environments answer completely different questions: Is there enough memory? What's the latency at 100 concurrent requests? Which GPU is most cost-effective? Will quantization drop quality? This article provides the complete deployment chain: fix metrics first, then calculate memory, then select quantization and batching strategies, and finally land on a runnable vLLM example. Underlying mechanisms (autoregressive loop, KV cache, sampling) are at Inference Fundamentals.

I. Fix Metrics First: Deployment Language Is Numbers ​

The deployment team must speak the same metric language:

MetricFull Name / MeaningMeasures WhatTypical Scale (Reference)
TTFTTime To First Token, first-token latencyHow long the user "waits for the first response"Hundreds of ms to seconds
TPOTTime Per Output Token, time per generated tokenGeneration speedTens to hundreds of ms/token
Throughputtokens/s, output volume per unit timeOverall service capacityHundreds to thousands of tokens/s per GPU
ConcurrencyNumber of requests served simultaneouslyCapacityStrongly related to memory / batching
text
User experience conversion: TTFT < 1s is "acceptable"; the faster TPOT, the stronger the "typewriter feel."
Throughput × tokens per request ÷ concurrency ≈ number of GPUs you need (rough estimate).

Latency vs. throughput trade-off

Latency-prioritized (interactive dialogue) and throughput-prioritized (offline batch processing) are two completely different optimization directions: the former reduces batch size to preserve speed; the latter increases batch size to maximize throughput. Clarify which one your business belongs to before tuning parameters.

Reference Ranges for Metric Targets and How to Set Them ​

Metrics don't have absolute "standard values," but reference ranges can be given by business type (actual values depend on testing and product requirements):

Business TypeTTFT TargetTPOT TargetPrimary Constraint
Real-time dialogue (Chat)< 1sAs fast as possible (typewriter feel)Latency-first
Customer service / ticket assistance< 2s< 100ms/tokenLatency + quality
Code completion< 500ms< 50ms/tokenExtreme latency sensitivity
Offline batch (summarization / extraction)Not sensitiveNot sensitiveThroughput + cost-first

Principle for setting targets: first measure the current baseline (P50 and P95), then set an "acceptable P95" rather than targeting the "average" — P95 represents the real user experience; averages are pulled down by a few fast requests.

II. Memory Estimation: Weights + KV Cache + Activations ​

Inference memory consists of three parts — learn to calculate by hand first, then trust any monitoring tool:

Total Memory ≈ Weight Memory + KV Cache Memory + Activation / Temporary Memory

1. Weight Memory ​

Weight Memory = Parameter Count × Bytes per Parameter
fp16/bf16: 2 bytes/parameter     INT8: 1 byte/parameter     INT4: 0.5 bytes/parameter
ModelParameter CountFP16INT8INT4
7B/8B~8B~16 GB~8 GB~4 GB
13B~13B~26 GB~13 GB~6.5 GB
70B~70B~140 GB~70 GB~35 GB

2. KV Cache Memory ​

KV cache stores the Key/Value generated at each layer for each sequence (mechanism at Transformer Architecture Explained), and grows linearly with: concurrent requests × context length:

KV memory per token = 2 × num_layers × num_kv_heads × head_dim × bytes_per_elem
KV cache total = KV memory per token × context_length × concurrency

Using Llama-2 70B (80 layers, 8 KV heads, head dim 128, fp16) as an example:

Per token = 2 × 80 × 8 × 128 × 2 bytes = 327,680 bytes ≈ 0.31 MB/token
4096 context × 32 concurrency ≈ 0.31MB × 4096 × 32 ≈ 40 GB !

These numbers reveal the first law of deployment: KV cache grows much faster than weights; with long context + high concurrency, it becomes the dominant memory consumer (this is an area that neither MoE sparse expert models nor quantization can fully address). Engineering optimizations for KV cache are in Section IV.

3. Activation Memory ​

Intermediate activations at each layer during the prefill phase also consume memory, proportional to batch size and sequence length, typically significantly compressed by fused operators like FlashAttention. For estimation, multiply the "weights + KV cache" result by 1.1–1.3 to leave activation headroom.

The difference between inference and training memory

Training memory also includes gradients and optimizer states (full-parameter training ≈ 12–16 bytes per parameter), 4–8× that of inference. "70B model, 140 GB memory" is the inference metric; training is calculated separately. Training-side estimation is in the QLoRA table at Fine-Tuning in Practice.

III. Quantization: Trading Precision for Memory and Speed ​

Quantization compresses weights from fp16 to low bit-width — the primary method for "running large models on small cards" and "reducing costs."

MethodPrecisionCharacteristicsTypical Tools
Dynamic quantization (INT8)8-bitGeneral-purpose, simple, minimal quality dropvLLM built-in, bitsandbytes
GPTQ4-bitPost-training quantization, dequantizes at activation time; significant memory savingsauto-gptq, supported by vLLM
AWQ4-bitActivation-aware quantization, smaller precision lossSupported by vLLM, autoawq
GGUF (K-quant)1–8-bitllama.cpp ecosystem, runs on CPU / edgellama.cpp, Ollama
FP88-bitNative support on newer GPUs (H series)TensorRT-LLM, vLLM
bash
# Load a quantized model directly with vLLM (supports both AWQ and GPTQ)
pip install vllm

# Load 4-bit AWQ version of Qwen2.5-7B
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-7B-Instruct-AWQ \
  --quantization awq \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192

Quantization is not free

4-bit quantization typically causes a 1–3% drop in benchmark scores (varies by task and model), but the inference speed and memory gains are enormous. Decision sequence: run fp16 baseline first, then run INT8/AWQ comparison, confirm the quality drop is acceptable using Evaluation in Practice methods, then launch.

IV. Batching and KV Cache Optimization ​

1. Continuous Batching: The Biggest Throughput Lever ​

Traditional batching is "static" (the next batch starts only after all current requests complete), causing short requests to wait for long ones. Continuous batching lets sequences enter and exit the batch at the token level — as soon as one completes, space is freed for a new request, significantly improving GPU utilization.

2. PagedAttention: KV Cache Memory Management ​

vLLM's core innovation, PagedAttention, manages KV cache in "pages" (analogous to virtual memory in operating systems), eliminating fragmentation and allocating on demand, increasing throughput 2–4× under the same memory (official benchmarks; actual numbers vary by workload, see References).

3. Quick Overview of Other Optimizations ​

OptimizationEffect
FlashAttentionFused operator, saves memory + speeds up
Prefix cachingReuse KV cache for identical system prompts
Speculative decodingSmall model drafts + large model verifies, reduces TPOT
Structured output constrained decodingSkip invalid tokens when generating JSON, saves time and stabilizes format

4. Multi-GPU Deployment: Tensor Parallelism ​

When a model doesn't fit on a single card, use Tensor Parallelism (TP) to split a layer across multiple GPUs, with GPUs synchronizing via high-speed interconnects (NVLink / InfiniBand). Key points:

Parallelism TypeWhat Gets SplitWhen to Use
Tensor parallelismMatrices / attention heads within a layer split across GPUsWhen the model doesn't fit on one card; sensitive to inter-GPU bandwidth
Pipeline parallelismLayers grouped to different GPUsMany layers; slight throughput decrease
Data parallelismEach GPU gets one model copy, data is splitThroughput scaling (used with TP)

vLLM enables TP with one line: --tensor-parallel-size 2; note that TP requires high-speed inter-GPU connectivity, otherwise communication will tank performance. 70B fp16 commonly uses 2×80GB or 4×40GB configurations.

V. GPU Selection ​

GPUMemoryReference Positioning (based on 2024–2025 mainstream market; prices fluctuate)
RTX 409024 GBConsumer grade: 7B at full precision / 13B INT4 on one card
L40S48 GB13B fp16 / 70B INT4 on one card
A100 40/80 GB40/80 GBDatacenter workhorse: 70B INT4 (80 GB)
H100 80 GB80 GBFlagship for large model training / inference, native FP8
H20 / domestic GPUsVariesChoice under constrained environments; check ecosystem compatibility

Selection decision table (how to pick a GPU):

Your Model + PrecisionMinimum Single-Card RequirementNotes
7B INT48–16 GB is enoughEdge / low-budget
7B fp1624 GBRTX 4090 class
13B fp16 / 70B INT440–48 GBL40S / A100
70B fp1680 GB × 2 (tensor parallelism)H100 / A100
High concurrencyIncrease memory per "KV cache formula"Don't just look at weights

Quick criterion for "is memory enough?"

Weight memory + KV cache (at peak concurrency × peak context) ≤ single-card memory × gpu-memory-utilization. If it exceeds, quantize first; if still not enough, use multi-GPU tensor parallelism or reduce concurrency. Always account for KV cache — accounting only for weights is the #1 source of deployment accidents.

VI. vLLM Deployment Example: End-to-End ​

vLLM provides an OpenAI-compatible API, so frontend code doesn't need a single line of change.

bash
# ① Start the service (using a 7B model as example)
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-7B-Instruct \
  --served-model-name my-llm \
  --gpu-memory-utilization 0.9 \
  --max-model-len 8192 \
  --enforce-eager          # Can disable CUDA graph for environments with poor compatibility

# ② Call via OpenAI SDK (client-side is transparent)
python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

resp = client.chat.completions.create(
    model="my-llm",
    messages=[{"role": "user", "content": "What is a KV cache?"}],
    temperature=0.7,
    max_tokens=512,
    stream=True,                       # Streaming output
)
for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="")
bash
# ③ Monitoring and load testing
curl http://localhost:8000/v1/models                     # View model list
# Load testing: use openai-benchmark / h2o-llmstudio or custom concurrency scripts
# Key monitored metrics: TTFT percentiles, TPOT percentiles, GPU utilization, memory usage, queue backlog

OpenAI-Compatible API: Protocol Details ​

vLLM's /v1/chat/completions is compatible with the OpenAI protocol. Four details that commonly trip people up:

DetailDescription
Streaming (stream)stream=true returns data: fragments (SSE), ending with data: [DONE]
Tool callingtools parameter and tool_calls response fields require client-side support
Structured outputJSON mode / response_format support varies by engine and model; test first
Auth and exposureSelf-hosted services need their own auth (API keys); don't expose publicly without protection
python
# SSE handling for streaming responses (client-side)
import requests

resp = requests.post(
    "http://localhost:8000/v1/chat/completions",
    json={"model": "my-llm",
          "messages": [{"role": "user", "content": "Hello"}],
          "stream": True},
    headers={"Authorization": "Bearer EMPTY"},
    stream=True,
)
for line in resp.iter_lines():
    if line and line.startswith(b"data: "):
        payload = line[6:]
        if payload == b"[DONE]":
            break
        # Parse JSON, take choices[0].delta.content and print fragment by fragment

Compatible ≠ functionally equivalent

"OpenAI compatible" is protocol-compatible, not functionally equivalent: different engines support extended fields like tools, response_format, and logprobs to different degrees. Test every field you plan to use before launch rather than assuming they're available.

Deployment Checklist (Verify Each Item Before Launch) ​

ItemCheck
Memory budgetWeights + KV cache (peak concurrency × peak context) ≤ available memory
Quantization validationINT4 vs. fp16 baseline difference has been evaluated
Metrics metTTFT/TPOT satisfy product requirements (including P95)
CompatibilityOpenAI-compatible API has been integrated (streaming / tool calling / JSON mode)
MonitoringLatency, throughput, memory, error rate, queue backlog all connected to alerts
RollbackModel/config versions are versioned and can be rolled back in one click

VII. Canary Deployment and Monitoring ​

Launch is not the end — it's the start of a new round of engineering:

MethodPractice
CanaryNew model serves only 5% of traffic; observe metrics and user feedback before scaling
Shadow modeOld and new models run in parallel; new model's results are logged but not shown, compared offline
MonitoringTTFT/TPOT P50/P95, error rate, queue length, GPU utilization
Evaluation linkageOnline spot-check results flow back into the golden set, forming a closed loop (see Evaluation in Practice)

Common deployment accident areas

Memory OOM (only counted weights, not KV cache), cascading failures from unlimited retries (set timeouts and backoff), streaming interruption without reconnection, difficult model version rollback. Write "how to fail gracefully" into the design, don't wait for it to happen.

Monitoring Metrics System: What to Watch After Launch ​

CategoryMetricAlert Threshold (Example)
LatencyP50, P95 of TTFT / TPOTP95 TTFT exceeds 2s
CapacityConcurrent requests, queue backlogQueue consistently backing up
ComputeGPU utilization, memory usageMemory usage > 95%
QualityError rate, retry rate, timeout rateError rate > 1%
BusinessCompletion rate, negative feedback rateNegative feedback rate increasing

The value of monitoring data lies in trends: memory usage slowly creeping from 60% to 90% signals traffic growth or a memory leak — something worth addressing before "the system crashes today."

VIII. Cloud vs. Local vs. API: Three Options for Decision ​

DimensionSelf-hosted GPUs (cloud/local)Managed API (OpenAI/Anthropic, etc.)Self-hosted open-source
Cost curveHigh fixed cost, low marginal costPay-per-use, no sunk costSame, but requires ops
Data securityControlled data egressData leaves network (evaluate compliance)Data stays entirely internal
Capability ceilingLimited by open-source modelsClosed-source flagships are strongestLimited by open-source models
Ops burdenHigh (K8s, GPU, elasticity)ZeroMedium–high
CustomizationFully controllableLimited (fine-tuning needs platform support)Fully controllable

Recommended decision sequence:

Data can leave network, need strongest capability, don't want ops → Managed API
Data-sensitive / compliance restrictions / deep customization → Self-host open-source (vLLM)
Large and stable volume → Self-hosted GPUs (or hybrid: API fallback + self-hosted primary)

Cost Accounting: A Simplified Example ​

Do the math before selection (numbers are for teaching only; prices change frequently; refer to official pricing):

Scenario: 7B model, 100K requests/day, avg input 1000 tokens, output 300 tokens
Daily token count ≈ 100K × (1000 + 300) ≈ 130M tokens

Option A: Managed API (assuming $0.004/M input tokens, $0.016/M output tokens)
  Cost ≈ 100K × (1000 × $0.004 + 300 × $0.016) / 1M ≈ $0.88/day ≈ $320/year

Option B: Self-host 1×RTX 4090 (~24 GB, can run 7B INT4)
  One-time hardware ~$400; electricity/ops extra; capacity ceiling ~tens to hundreds of thousands of tokens/min

Conclusion pattern: low volume → API (zero ops), high and stable volume → self-host (amortize fixed cost), middle ground → hybrid. Don't forget to include dev/ops labor cost — that's often the most expensive part.

Further Reading ​

References ​