Appearance
Deploying and Optimizing LLM Inference
The goal of deploying and optimizing LLM inference is to turn a chosen, fine-tuned LLM into a production service that is stable, affordable, and fast. The concepts page Inference Optimization and Quantization explains in depth why inference on GPUs is slow and how quantization and the KV cache work; this page tackles the practical side: cloud API or self-hosting, how to export and quantize the model, how to stand up the service, how to wire in the API, how to configure caching and rate limiting, and how to monitor in production — every step comes with a minimal configuration you can copy.
First, set realistic expectations: deploying an LLM is fundamentally different from running a traditional web service — it runs on GPUs, its output is probabilistic, and latency and throughput swing sharply with load. Every practice on this page therefore serves exactly three implicit goals: push latency down, push throughput up, and cut costs — and find the balance among the three that fits your business.
How to read this page
Start with Inference Optimization and Quantization to build a mechanistic understanding, then work through this page starting from "Choosing a Deployment Path". If you haven't picked a model yet, see Model & Leaderboard Quick Reference; if your model isn't fine-tuned yet, start with Fine-Tuning Your Own LLM. Once deployment is done, hand evaluation and regression testing to Setting Up LLM Evaluation.
1. Choosing a Deployment Path: Cloud API or Self-Hosting?
The first step isn't writing configuration — it's deciding where the model runs. The real difference between the two routes isn't "expensive vs. cheap" but your priorities across four dimensions: cost, control, latency, and data security.
| Decision factor | Cloud API (OpenAI, Anthropic, Chinese LLM platforms, etc.) | Self-hosting (vLLM / Ollama / TensorRT-LLM) |
|---|---|---|
| Upfront cost | Nearly zero; pay per token | High: GPU hardware or rented cloud GPUs + ops headcount |
| Marginal cost | Every call costs money; the more you use, the more you pay | Fixed cost; once capacity is saturated, marginal cost approaches zero |
| Latency | Network hops + queueing, fairly volatile (shared compute) | Direct intranet access, predictable; dedicated compute can be optimized to the extreme |
| Control | Low: can't tune sampling details, can't instrument, upgrades at the provider's mercy | High: decoding parameters, batching, prefix caching all under your control |
| Data security | Data egress and retention policies dictated by the provider | Data never leaves your intranet; satisfies private-deployment/compliance requirements |
| Elasticity | Naturally elastic; absorbs traffic spikes | You scale it yourself; capacity planning is routine work |
| Ops burden | Borne by the provider | Full-stack self-management: GPUs, drivers, frameworks, version upgrades |
| Customization | Limited to the vendor's models | You can deploy your own fine-tuned models |
The quick call (rule of thumb):
- Small daily request volume, no plans to change models, no desire to own GPUs → cloud API, and spend the engineering effort on the product instead;
- Volume high enough that "the monthly cloud API bill exceeds a GPU server", or the model must stay private → self-hosting;
- The most economical common setup is hybrid: self-hosted open-source models by default, falling back to a cloud API for peak traffic or unusually hard requests.
text
Decision flow (start at step 1; take the first route that matches):
1. Business requires data to stay inside the network / pass compliance audits? → Self-hosting (no choice)
2. Need to deploy your own fine-tuned model (LoRA etc.)? → Self-hosting
3. Projected monthly volume under some token threshold (say, a few hundred million)? → Cloud API
4. Extremely strict latency requirements (e.g., first token < 500ms) at large volume? → Self-hosting
5. None of the above? → Hybrid: self-host by default + cloud API at peakDon't be fooled by "free and open source"
Open-source models are free, but the GPUs that run them are not. The true cost of self-hosting = hardware/rental + electricity + ops headcount, and the combined hidden cost of those three is routinely underestimated. For model-selection trade-offs, see Model & Leaderboard Quick Reference.
2. Step-by-Step: From Model Files to a Production Service
The walkthrough below uses the self-hosting route as the main thread (for the cloud API route, simply swap the "server side" for a vendor endpoint), covering five steps: export & quantize → serve with vLLM → OpenAI-compatible API integration → gateway & caching → monitoring & alerting.
Step 1: Model Export and Quantization
Before self-hosting, export the model weights to a standard format (safetensors), then quantize according to your VRAM budget. Quantization is self-hosting's highest-leverage tool: on the same VRAM, 4-bit quantization nearly doubles the size of model you can deploy. For the underlying principles, see the "Quantization" section of Inference Optimization and Quantization.
Comparison of mainstream quantization options (as of mid-2025):
| Option | Type | Characteristics | Best for |
|---|---|---|---|
| AWQ | 4-bit weight quantization (activation-aware) | Needs little calibration data, small accuracy loss, native vLLM support | The sensible default for general production |
| GPTQ | 4-bit weight quantization (Hessian-based) | Long-established option with a mature ecosystem | Compatibility with older toolchains |
| FP8 | 8-bit floating point | Hardware-accelerated on newer cards like the H100 | When VRAM is plentiful and precision matters |
| GGUF (Q4_K_M, etc.) | 4-bit quantization | Ollama / llama.cpp ecosystem | Personal machines, edge devices |
For example, quantize and export with AWQ using AutoAWQ:
python
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer
model_path = "Qwen/Qwen2.5-7B-Instruct" # or the path to your fine-tuned model
quant_path = "./qwen2.5-7b-awq" # output directory for the quantized model
model = AutoAWQForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path)
# AWQ needs a small calibration set (128~512 samples is enough; use representative samples from your target task)
calib_data = [
"What is the company's refund policy?",
"Please summarize the core clauses of this contract.",
# ... about 128~512 samples in total
]
model.quantize(tokenizer, quant_config={"zero_point": True, "q_group_size": 128, "w_bit": 4})
model.save_quantized(quant_path)
tokenizer.save_pretrained(quant_path)Think through three things before quantizing
- Quantization is irreversible: archive the original weights — a quantized model cannot be losslessly restored;
- Validate the accuracy loss: run the same eval set before and after quantization and compare; if the drop exceeds your threshold (e.g., >1%), consider a wider bit width — see LLM Evaluation and Benchmarks;
- Quantize the inference side first: FP16 for training/fine-tuning, INT4 for deployment — two sets of weights, each doing its own job.
Step 2: Serving with vLLM
vLLM is the de facto standard for self-hosting LLMs today: it provides PagedAttention-based KV cache management, continuous batching, and an OpenAI-compatible API, with throughput typically an order of magnitude higher than plain Transformers inference. Minimal launch command:
bash
# Install (requires a CUDA environment; see the official docs)
pip install vllm
# Start an OpenAI-compatible server (reserve VRAM per GPU, cap the output length)
vllm serve Qwen/Qwen2.5-7B-Instruct \
--quantization awq \
--gpu-memory-utilization 0.90 \
--max-model-len 8192 \
--port 8000More elaborate production settings go into launch parameters or environment variables:
| Parameter | Recommended value | Notes |
|---|---|---|
--gpu-memory-utilization | 0.85~0.95 | GPU memory utilization; too high risks OOM, too low wastes VRAM |
--max-model-len | Set per your business (e.g., 8192) | Determines maximum context length; note it also constrains the KV cache reservation |
--max-num-seqs | 64~256 | Maximum sequences per batch; higher raises throughput but lifts latency |
--tensor-parallel-size | Number of GPUs (for multi-GPU) | Tensor parallelism degree |
--enable-prefix-caching | Enabled | Prefix caching; hit rates are substantial in RAG workloads |
--served-model-name | Custom | Model name exposed to clients; can be an alias for the real path |
yaml
# docker-compose.yaml example: single-GPU deployment + health check
services:
llm:
image: vllm/vllm-openai:latest
command:
- "--model"
- "/models/qwen2.5-7b-awq"
- "--quantization"
- "awq"
- "--gpu-memory-utilization"
- "0.90"
- "--max-model-len"
- "8192"
- "--enable-prefix-caching"
volumes:
- ./models:/models
ports:
- "8000:8000"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]Smoke-test after startup: curl http://localhost:8000/v1/models should return the model list; /health or /v1/chat/completions works as a liveness probe.
Step 3: OpenAI-Compatible API Integration
vLLM exposes an OpenAI-compatible interface, which means client code is nearly identical to what you'd write against a cloud API — you just switch endpoints. This also makes "hybrid deployment" (self-hosted with cloud API fallback) nearly free to implement.
python
from openai import OpenAI
# Only base_url and api_key change; everything else is identical to calling OpenAI's cloud
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="not-needed", # vLLM does not validate it by default
timeout=60.0,
)
resp = client.chat.completions.create(
model="qwen2.5-7b-awq",
messages=[
{"role": "system", "content": "You are a knowledge-base Q&A assistant."},
{"role": "user", "content": "What is the company's refund policy?"},
],
temperature=0.2,
max_tokens=512,
stream=False,
)
print(resp.choices[0].message.content)Streaming output (lower first-token latency, better conversational experience):
python
stream = client.chat.completions.create(
model="qwen2.5-7b-awq",
messages=messages,
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")The integration layer's "three iron rules"
- Always set timeout and max_tokens: so a hung or runaway model can't drag down your entire call chain;
- Design streaming and timeout retries separately: retrying after a timeout in streaming scenarios easily produces duplicated output; auto-retry only for non-streaming + idempotent calls;
- Refer to models by alias:
--served-model-namedecouples the application from the underlying model, so you can swap models without touching application code.
Step 4: Gateway and Caching
The service is up; next comes the gateway layer that both protects it and squeezes every bit of performance out of it. Two things are mandatory: semantic/exact caching (to cut down repeated computation) and rate limiting (so no single client can saturate the service).
Redis caching: for exactly identical requests, return the previous result directly and skip a lot of duplicate inference. The more advanced semantic cache matches similar questions by vector similarity (see Vector Databases and Semantic Search for how vector retrieval works) — a higher hit rate, but you take on the risk of false recalls.
python
import hashlib
import redis
r = redis.Redis(host="localhost", port=6379, decode_responses=True)
def cached_chat(messages, ttl=3600):
# Exact cache: hit directly on "the same message sequence"
key = "llm:cache:" + hashlib.sha256(
repr(messages).encode("utf-8")
).hexdigest()
hit = r.get(key)
if hit:
return hit
resp = llm_client.chat.completions.create(
model=MODEL, messages=messages, temperature=0
)
text = resp.choices[0].message.content
r.setex(key, ttl, text) # TTL keeps the cache from growing without bound
return textRate limiting (gateway layer: a token bucket to cap QPS and concurrency):
yaml
# Nginx + lua-resty-limit-traffic example (or simply use a gateway product such as Kong/APISIX)
limit_req_zone $binary_remote_addr zone=llm_per_ip:10m rate=10r/s;
server {
location /v1/chat/completions {
limit_req zone=llm_per_ip burst=20 nodelay;
proxy_pass http://vllm-svc:8000;
proxy_read_timeout 300s; # long outputs call for a relaxed read timeout
}
}Three small things worth doing at the gateway layer
- Queueing and backpressure: when concurrency exceeds what the model can handle, queue requests instead of knocking the service over, and pair this with
max-num-seqsto cap model-side concurrency; - Failover switch: fall back to the cloud API automatically if the self-hosted service dies (just repoint
base_urlat the cloud endpoint); - Log every call's inputs, outputs, and metadata — this is both an audit trail and a data asset for later evaluation and prompt iteration (used by Setting Up LLM Evaluation).
Step 5: Monitoring and Alerting
An LLM service needs monitoring at three layers: system (GPU utilization, VRAM, temperature), inference metrics (TTFT / TPOT / tokens/s), and quality (whether output matches expectations). The inference metrics are LLM-specific; definitions below (detailed methodology in LLM Evaluation and Benchmarks):
| Metric | Full name / meaning | Unit | Notes |
|---|---|---|---|
| TTFT | Time To First Token, time to the first token | ms | The user-perceived "how fast does it respond"; lower is better |
| TPOT | Time Per Output Token, average time per output token | ms/token | The perceived "speed" during generation |
| tokens/s | Throughput: tokens generated per second | tok/s | How much capacity you have under batch load |
| QPS | Queries per second | req/s | Business load |
| KV cache utilization | Prefix cache hit rate / VRAM usage | % | Determines the throughput ceiling |
Wire vLLM's Prometheus metrics into your monitoring stack (vLLM ships with a /metrics endpoint exposing vllm:num_requests_running, vllm:time_to_first_token_seconds, vllm:generation_tokens_total, and more):
yaml
# prometheus.yml snippet
scrape_configs:
- job_name: vllm
metrics_path: /metrics
static_configs:
- targets: ["vllm-svc:8000"]yaml
# alertmanager alert-rule example (set thresholds on the 99th percentile; run a baseline first, then decide)
groups:
- name: llm
rules:
- alert: TTFTHigh
expr: histogram_quantile(0.99, vllm_time_to_first_token_seconds_bucket) > 2
for: 5m
annotations:
summary: "First-token latency above 2s — check GPU load and queue depth"
- alert: GPUOOMRisk
expr: (1 - vllm_gpu_cache_usage_perc) < 0.1
for: 3m
annotations:
summary: "KV cache usage above 90% — OOM risk"Monitoring: "baseline first, thresholds second"
Don't pick alert thresholds on gut feel. Run the service bare for a week to capture the normal fluctuation range of p99 TTFT and TPOT, then set alert lines — otherwise alerts will either fire falsely every day until people ignore them, or stay silent when something actually breaks. For the full methodology on thresholds and evaluation criteria, see LLM Evaluation and Benchmarks.
3. Performance Tuning Checklist: The Order of the Four Knobs
Performance optimization has four main knobs. Turn them in this order — the order follows "cheapest to change first, most certain payoff first":
| Order | Knob | How | Payoff and cost |
|---|---|---|---|
| 1 | Batching | Enable continuous batching (vLLM default), coalesce small requests; share prefixes in RAG scenarios | The most significant throughput gain at nearly zero cost; the price is a slight rise in per-request latency |
| 2 | KV cache optimization | Turn on --enable-prefix-caching; for multi-turn conversations use the chat endpoint instead of rebuilding history every time | Big throughput gains in long-context + many-user scenarios; nearly free |
| 3 | Concurrency and queueing | Tune --max-num-seqs, load-test to find the throughput/latency balance point; rate limiting at the gateway as protection | Needs load-test data to back it up, otherwise you'll fix one thing and break another |
| 4 | VRAM and quantization | Push --gpu-memory-utilization to a safe ceiling; quantize with AWQ/GPTQ if needed | Directly raises the capacity ceiling; quantization carries accuracy risk and needs validation |
The one-line verdict
Fix batching before touching quantization. 80% of throughput problems are solved simply by "merging requests and running them together" — quantization rarely even comes into play. It is the last resort because of its accuracy risk.
4. Cost Optimization: Make Every Dollar Produce Tokens
The heart of inference cost optimization is eliminating wasted computation. Four levers, in priority order:
- Cache hit rate: exact caching + semantic caching (see Step 4); in common FAQ scenarios hit rates can reach 30%~50%, cutting costs by the same fraction;
- Model routing: send simple requests to a small model and complex ones to a large model — the highest-ROI and most overlooked move. Use a small classifier or simple rules to judge "how hard is this query" first; only the hard ones go to the big model;
- Quantization: INT4 inference cuts cost and VRAM immediately (Inference Optimization and Quantization);
- Output length control: cap
max_tokens, add length constraints to prompts, truncate overlong outputs — you pay per generated token, so shrinking output length is a direct saving.
python
# Model routing skeleton: simple questions to the small model, complex ones to the large model
def route(query: str) -> str:
complexity = estimate_complexity(query) # rules / small classifier / heuristics
model = "qwen2.5-7b-awq" if complexity < 0.5 else "qwen2.5-32b-awq"
return call_llm(model, query)5. Security and Compliance: Deployment's Last Gate
The moment your model goes live, it presents an attack surface. At minimum, do these four things (full governance framework in AI Safety and Governance):
| Measure | How | Where it applies |
|---|---|---|
| Content filtering | Attach a moderation model/rules on the output side to block policy-violating content | Consumer-facing chat and writing products |
| Data masking | On the request side, detect and replace PII first (phone numbers, ID numbers, emails, etc.), then restore it after the response | Any scenario handling users' private data |
| Audit logging | Record inputs/outputs, model version, time, user identity; leave a trail for sensitive operations | Compliance audits, incident tracing |
| Permissions and isolation | For agent-style applications, restrict tool permissions and require human review of high-risk operations | Agent applications and tool-calling scenarios |
python
# Minimal masking example: replace sensitive data before the request hits the model, restore it after the response
import re
def mask_pii(text: str):
text = re.sub(r"1[3-9]\d{9}", "[phone number]", text) # phone number
text = re.sub(r"\d{17}[\dXx]", "[ID number]", text) # ID card number
return textSelf-hosting does not mean you're safe by default
Self-hosting only keeps data inside your intranet — content safety and prompt-injection defenses must not be skipped at all. Prompt injection, jailbreaks, and similar attacks work just as well against self-hosted models, and may even land more easily because the model is more "exposed". Security is a precondition for going live, not a patch applied afterward.
6. Common Pitfalls: Read This List Before You Ship
| Pitfall | Symptom | Prevention and fix |
|---|---|---|
| VRAM OOM | Crash at startup or CUDA out of memory at runtime | Quantize before launch; start gpu-memory-utilization at 0.85 and tune up; budget VRAM for the KV cache and max-model-len together |
| Timeouts | The client times out before the server does; users see "no response" | Streaming output lowers perceived first-token latency; relax the gateway's proxy_read_timeout; size client timeouts from max_tokens |
| Concurrency exhausted | One user's batch of requests occupies every concurrency slot while everyone else queues | Gateway rate limiting + per-user quotas; cap it with max-num-seqs; queueing and backpressure as the safety net |
| Runaway output length | Answers balloon endlessly, token bills explode, APIs time out | Add length constraints to prompts (see Prompt Engineering Playbook); enforce a hard max_tokens cap; truncate overlong outputs |
| Hallucinations reaching production | Answers confidently fabricate facts; user complaints | Ground with RAG + "answer only from the provided material" (Building a RAG App from Scratch); add output-side validation |
| Silent accuracy regression | Quality degrades after quantization; users churn | Compare on the same eval set before and after quantization (Setting Up LLM Evaluation) |
| No version management | After switching model/quantization versions, you can't roll back | Put model files and configs under version control; alias with served-model-name for one-click switching |
For a more complete list of "anti-patterns" (including evaluation leakage, data contamination, and other pitfalls beyond deployment), see Common Pitfalls and Anti-Patterns.
Further Reading
- Inference Optimization and Quantization — the mechanistic foundation for this page: why inference is slow, how quantization and the KV cache work
- Model & Leaderboard Quick Reference — model selection and cost trade-off data for self-hosting vs. cloud APIs
- Large Language Models (LLM) — the architecture and capability limits of what you're deploying
- Fine-Tuning Your Own LLM — upstream of deployment: where models come from
- Building a RAG App from Scratch — end-to-end deployment of a RAG application (retrieval service included)
- Setting Up LLM Evaluation — evaluation and regression testing before launch
- LLM Evaluation and Benchmarks — metric definitions and benchmarks beyond TTFT/TPOT
- AI Safety and Governance — the full picture of content filtering, prompt injection, and data compliance
- Prompt Engineering Playbook — post-launch prompt iteration and A/B testing
- AI Agents — deploying agent-style applications and isolating tool permissions
- Glossary — consistent definitions for deployment-related terms
References
- vLLM Official Documentation — the authoritative reference for deployment, performance tuning, Prometheus metrics, and quantization support
- vLLM: Engine Arguments (official parameter reference) — full explanations of
gpu-memory-utilization,max-model-len,max-num-seqs, and other parameters - Ollama Official Documentation — a lightweight option for GGUF quantization and local deployment
- AutoAWQ Official Repository — the AWQ quantization implementation and usage examples
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (paper) — the original GPTQ quantization algorithm paper
- Lin et al. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (MLSys 2024) — the original AWQ quantization algorithm paper
- OpenAI: Production best practices — best practices for latency, concurrency, and monitoring (cloud API route)
- NVIDIA: LLM Inference Performance Optimization Guide (TensorRT-LLM docs) — reference solutions for extreme-performance scenarios
- MLPerf Inference — inference performance benchmarks; see where your service stands among similar hardware