Skip to content

Deploying and Optimizing LLM Inference

At a glance Turn a trained LLM into a stable, inexpensive, fast production service, covering the cloud API vs. self-hosting decision, AWQ/GPTQ quantization export, vLLM launch configuration, OpenAI-compatible API integration, Redis caching and rate limiting, TTFT/TPOT monitoring and alerting, plus cost optimization, security and compliance, and troubleshooting common pitfalls.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Deploying and Optimizing LLM Inference ​

The goal of deploying and optimizing LLM inference is to turn a chosen, fine-tuned LLM into a production service that is stable, affordable, and fast. The concepts page Inference Optimization and Quantization explains in depth why inference on GPUs is slow and how quantization and the KV cache work; this page tackles the practical side: cloud API or self-hosting, how to export and quantize the model, how to stand up the service, how to wire in the API, how to configure caching and rate limiting, and how to monitor in production — every step comes with a minimal configuration you can copy.

First, set realistic expectations: deploying an LLM is fundamentally different from running a traditional web service — it runs on GPUs, its output is probabilistic, and latency and throughput swing sharply with load. Every practice on this page therefore serves exactly three implicit goals: push latency down, push throughput up, and cut costs — and find the balance among the three that fits your business.

How to read this page

Start with Inference Optimization and Quantization to build a mechanistic understanding, then work through this page starting from "Choosing a Deployment Path". If you haven't picked a model yet, see Model & Leaderboard Quick Reference; if your model isn't fine-tuned yet, start with Fine-Tuning Your Own LLM. Once deployment is done, hand evaluation and regression testing to Setting Up LLM Evaluation.

1. Choosing a Deployment Path: Cloud API or Self-Hosting? ​

The first step isn't writing configuration — it's deciding where the model runs. The real difference between the two routes isn't "expensive vs. cheap" but your priorities across four dimensions: cost, control, latency, and data security.

Decision factorCloud API (OpenAI, Anthropic, Chinese LLM platforms, etc.)Self-hosting (vLLM / Ollama / TensorRT-LLM)
Upfront costNearly zero; pay per tokenHigh: GPU hardware or rented cloud GPUs + ops headcount
Marginal costEvery call costs money; the more you use, the more you payFixed cost; once capacity is saturated, marginal cost approaches zero
LatencyNetwork hops + queueing, fairly volatile (shared compute)Direct intranet access, predictable; dedicated compute can be optimized to the extreme
ControlLow: can't tune sampling details, can't instrument, upgrades at the provider's mercyHigh: decoding parameters, batching, prefix caching all under your control
Data securityData egress and retention policies dictated by the providerData never leaves your intranet; satisfies private-deployment/compliance requirements
ElasticityNaturally elastic; absorbs traffic spikesYou scale it yourself; capacity planning is routine work
Ops burdenBorne by the providerFull-stack self-management: GPUs, drivers, frameworks, version upgrades
CustomizationLimited to the vendor's modelsYou can deploy your own fine-tuned models

The quick call (rule of thumb):

  • Small daily request volume, no plans to change models, no desire to own GPUs → cloud API, and spend the engineering effort on the product instead;
  • Volume high enough that "the monthly cloud API bill exceeds a GPU server", or the model must stay private → self-hosting;
  • The most economical common setup is hybrid: self-hosted open-source models by default, falling back to a cloud API for peak traffic or unusually hard requests.
text
Decision flow (start at step 1; take the first route that matches):
1. Business requires data to stay inside the network / pass compliance audits? → Self-hosting (no choice)
2. Need to deploy your own fine-tuned model (LoRA etc.)? → Self-hosting
3. Projected monthly volume under some token threshold (say, a few hundred million)? → Cloud API
4. Extremely strict latency requirements (e.g., first token < 500ms) at large volume? → Self-hosting
5. None of the above? → Hybrid: self-host by default + cloud API at peak

Don't be fooled by "free and open source"

Open-source models are free, but the GPUs that run them are not. The true cost of self-hosting = hardware/rental + electricity + ops headcount, and the combined hidden cost of those three is routinely underestimated. For model-selection trade-offs, see Model & Leaderboard Quick Reference.

2. Step-by-Step: From Model Files to a Production Service ​

The walkthrough below uses the self-hosting route as the main thread (for the cloud API route, simply swap the "server side" for a vendor endpoint), covering five steps: export & quantize → serve with vLLM → OpenAI-compatible API integration → gateway & caching → monitoring & alerting.

Step 1: Model Export and Quantization ​

Before self-hosting, export the model weights to a standard format (safetensors), then quantize according to your VRAM budget. Quantization is self-hosting's highest-leverage tool: on the same VRAM, 4-bit quantization nearly doubles the size of model you can deploy. For the underlying principles, see the "Quantization" section of Inference Optimization and Quantization.

Comparison of mainstream quantization options (as of mid-2025):

OptionTypeCharacteristicsBest for
AWQ4-bit weight quantization (activation-aware)Needs little calibration data, small accuracy loss, native vLLM supportThe sensible default for general production
GPTQ4-bit weight quantization (Hessian-based)Long-established option with a mature ecosystemCompatibility with older toolchains
FP88-bit floating pointHardware-accelerated on newer cards like the H100When VRAM is plentiful and precision matters
GGUF (Q4_K_M, etc.)4-bit quantizationOllama / llama.cpp ecosystemPersonal machines, edge devices

For example, quantize and export with AWQ using AutoAWQ:

python
from awq import AutoAWQForCausalLM
from transformers import AutoTokenizer

model_path = "Qwen/Qwen2.5-7B-Instruct"      # or the path to your fine-tuned model
quant_path = "./qwen2.5-7b-awq"              # output directory for the quantized model

model = AutoAWQForCausalLM.from_pretrained(model_path)
tokenizer = AutoTokenizer.from_pretrained(model_path)

# AWQ needs a small calibration set (128~512 samples is enough; use representative samples from your target task)
calib_data = [
    "What is the company's refund policy?",
    "Please summarize the core clauses of this contract.",
    # ... about 128~512 samples in total
]

model.quantize(tokenizer, quant_config={"zero_point": True, "q_group_size": 128, "w_bit": 4})
model.save_quantized(quant_path)
tokenizer.save_pretrained(quant_path)

Think through three things before quantizing

  1. Quantization is irreversible: archive the original weights — a quantized model cannot be losslessly restored;
  2. Validate the accuracy loss: run the same eval set before and after quantization and compare; if the drop exceeds your threshold (e.g., >1%), consider a wider bit width — see LLM Evaluation and Benchmarks;
  3. Quantize the inference side first: FP16 for training/fine-tuning, INT4 for deployment — two sets of weights, each doing its own job.

Step 2: Serving with vLLM ​

vLLM is the de facto standard for self-hosting LLMs today: it provides PagedAttention-based KV cache management, continuous batching, and an OpenAI-compatible API, with throughput typically an order of magnitude higher than plain Transformers inference. Minimal launch command:

bash
# Install (requires a CUDA environment; see the official docs)
pip install vllm

# Start an OpenAI-compatible server (reserve VRAM per GPU, cap the output length)
vllm serve Qwen/Qwen2.5-7B-Instruct \
  --quantization awq \
  --gpu-memory-utilization 0.90 \
  --max-model-len 8192 \
  --port 8000

More elaborate production settings go into launch parameters or environment variables:

ParameterRecommended valueNotes
--gpu-memory-utilization0.85~0.95GPU memory utilization; too high risks OOM, too low wastes VRAM
--max-model-lenSet per your business (e.g., 8192)Determines maximum context length; note it also constrains the KV cache reservation
--max-num-seqs64~256Maximum sequences per batch; higher raises throughput but lifts latency
--tensor-parallel-sizeNumber of GPUs (for multi-GPU)Tensor parallelism degree
--enable-prefix-cachingEnabledPrefix caching; hit rates are substantial in RAG workloads
--served-model-nameCustomModel name exposed to clients; can be an alias for the real path
yaml
# docker-compose.yaml example: single-GPU deployment + health check
services:
  llm:
    image: vllm/vllm-openai:latest
    command:
      - "--model"
      - "/models/qwen2.5-7b-awq"
      - "--quantization"
      - "awq"
      - "--gpu-memory-utilization"
      - "0.90"
      - "--max-model-len"
      - "8192"
      - "--enable-prefix-caching"
    volumes:
      - ./models:/models
    ports:
      - "8000:8000"
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: 1
              capabilities: [gpu]

Smoke-test after startup: curl http://localhost:8000/v1/models should return the model list; /health or /v1/chat/completions works as a liveness probe.

Step 3: OpenAI-Compatible API Integration ​

vLLM exposes an OpenAI-compatible interface, which means client code is nearly identical to what you'd write against a cloud API — you just switch endpoints. This also makes "hybrid deployment" (self-hosted with cloud API fallback) nearly free to implement.

python
from openai import OpenAI

# Only base_url and api_key change; everything else is identical to calling OpenAI's cloud
client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed",        # vLLM does not validate it by default
    timeout=60.0,
)

resp = client.chat.completions.create(
    model="qwen2.5-7b-awq",
    messages=[
        {"role": "system", "content": "You are a knowledge-base Q&A assistant."},
        {"role": "user", "content": "What is the company's refund policy?"},
    ],
    temperature=0.2,
    max_tokens=512,
    stream=False,
)
print(resp.choices[0].message.content)

Streaming output (lower first-token latency, better conversational experience):

python
stream = client.chat.completions.create(
    model="qwen2.5-7b-awq",
    messages=messages,
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

The integration layer's "three iron rules"

  1. Always set timeout and max_tokens: so a hung or runaway model can't drag down your entire call chain;
  2. Design streaming and timeout retries separately: retrying after a timeout in streaming scenarios easily produces duplicated output; auto-retry only for non-streaming + idempotent calls;
  3. Refer to models by alias: --served-model-name decouples the application from the underlying model, so you can swap models without touching application code.

Step 4: Gateway and Caching ​

The service is up; next comes the gateway layer that both protects it and squeezes every bit of performance out of it. Two things are mandatory: semantic/exact caching (to cut down repeated computation) and rate limiting (so no single client can saturate the service).

Redis caching: for exactly identical requests, return the previous result directly and skip a lot of duplicate inference. The more advanced semantic cache matches similar questions by vector similarity (see Vector Databases and Semantic Search for how vector retrieval works) — a higher hit rate, but you take on the risk of false recalls.

python
import hashlib
import redis

r = redis.Redis(host="localhost", port=6379, decode_responses=True)

def cached_chat(messages, ttl=3600):
    # Exact cache: hit directly on "the same message sequence"
    key = "llm:cache:" + hashlib.sha256(
        repr(messages).encode("utf-8")
    ).hexdigest()
    hit = r.get(key)
    if hit:
        return hit
    resp = llm_client.chat.completions.create(
        model=MODEL, messages=messages, temperature=0
    )
    text = resp.choices[0].message.content
    r.setex(key, ttl, text)        # TTL keeps the cache from growing without bound
    return text

Rate limiting (gateway layer: a token bucket to cap QPS and concurrency):

yaml
# Nginx + lua-resty-limit-traffic example (or simply use a gateway product such as Kong/APISIX)
limit_req_zone $binary_remote_addr zone=llm_per_ip:10m rate=10r/s;
server {
    location /v1/chat/completions {
        limit_req zone=llm_per_ip burst=20 nodelay;
        proxy_pass http://vllm-svc:8000;
        proxy_read_timeout 300s;   # long outputs call for a relaxed read timeout
    }
}

Three small things worth doing at the gateway layer

  1. Queueing and backpressure: when concurrency exceeds what the model can handle, queue requests instead of knocking the service over, and pair this with max-num-seqs to cap model-side concurrency;
  2. Failover switch: fall back to the cloud API automatically if the self-hosted service dies (just repoint base_url at the cloud endpoint);
  3. Log every call's inputs, outputs, and metadata — this is both an audit trail and a data asset for later evaluation and prompt iteration (used by Setting Up LLM Evaluation).

Step 5: Monitoring and Alerting ​

An LLM service needs monitoring at three layers: system (GPU utilization, VRAM, temperature), inference metrics (TTFT / TPOT / tokens/s), and quality (whether output matches expectations). The inference metrics are LLM-specific; definitions below (detailed methodology in LLM Evaluation and Benchmarks):

MetricFull name / meaningUnitNotes
TTFTTime To First Token, time to the first tokenmsThe user-perceived "how fast does it respond"; lower is better
TPOTTime Per Output Token, average time per output tokenms/tokenThe perceived "speed" during generation
tokens/sThroughput: tokens generated per secondtok/sHow much capacity you have under batch load
QPSQueries per secondreq/sBusiness load
KV cache utilizationPrefix cache hit rate / VRAM usage%Determines the throughput ceiling

Wire vLLM's Prometheus metrics into your monitoring stack (vLLM ships with a /metrics endpoint exposing vllm:num_requests_running, vllm:time_to_first_token_seconds, vllm:generation_tokens_total, and more):

yaml
# prometheus.yml snippet
scrape_configs:
  - job_name: vllm
    metrics_path: /metrics
    static_configs:
      - targets: ["vllm-svc:8000"]
yaml
# alertmanager alert-rule example (set thresholds on the 99th percentile; run a baseline first, then decide)
groups:
  - name: llm
    rules:
      - alert: TTFTHigh
        expr: histogram_quantile(0.99, vllm_time_to_first_token_seconds_bucket) > 2
        for: 5m
        annotations:
          summary: "First-token latency above 2s — check GPU load and queue depth"
      - alert: GPUOOMRisk
        expr: (1 - vllm_gpu_cache_usage_perc) < 0.1
        for: 3m
        annotations:
          summary: "KV cache usage above 90% — OOM risk"

Monitoring: "baseline first, thresholds second"

Don't pick alert thresholds on gut feel. Run the service bare for a week to capture the normal fluctuation range of p99 TTFT and TPOT, then set alert lines — otherwise alerts will either fire falsely every day until people ignore them, or stay silent when something actually breaks. For the full methodology on thresholds and evaluation criteria, see LLM Evaluation and Benchmarks.

3. Performance Tuning Checklist: The Order of the Four Knobs ​

Performance optimization has four main knobs. Turn them in this order — the order follows "cheapest to change first, most certain payoff first":

OrderKnobHowPayoff and cost
1BatchingEnable continuous batching (vLLM default), coalesce small requests; share prefixes in RAG scenariosThe most significant throughput gain at nearly zero cost; the price is a slight rise in per-request latency
2KV cache optimizationTurn on --enable-prefix-caching; for multi-turn conversations use the chat endpoint instead of rebuilding history every timeBig throughput gains in long-context + many-user scenarios; nearly free
3Concurrency and queueingTune --max-num-seqs, load-test to find the throughput/latency balance point; rate limiting at the gateway as protectionNeeds load-test data to back it up, otherwise you'll fix one thing and break another
4VRAM and quantizationPush --gpu-memory-utilization to a safe ceiling; quantize with AWQ/GPTQ if neededDirectly raises the capacity ceiling; quantization carries accuracy risk and needs validation

The one-line verdict

Fix batching before touching quantization. 80% of throughput problems are solved simply by "merging requests and running them together" — quantization rarely even comes into play. It is the last resort because of its accuracy risk.

4. Cost Optimization: Make Every Dollar Produce Tokens ​

The heart of inference cost optimization is eliminating wasted computation. Four levers, in priority order:

  1. Cache hit rate: exact caching + semantic caching (see Step 4); in common FAQ scenarios hit rates can reach 30%~50%, cutting costs by the same fraction;
  2. Model routing: send simple requests to a small model and complex ones to a large model — the highest-ROI and most overlooked move. Use a small classifier or simple rules to judge "how hard is this query" first; only the hard ones go to the big model;
  3. Quantization: INT4 inference cuts cost and VRAM immediately (Inference Optimization and Quantization);
  4. Output length control: cap max_tokens, add length constraints to prompts, truncate overlong outputs — you pay per generated token, so shrinking output length is a direct saving.
python
# Model routing skeleton: simple questions to the small model, complex ones to the large model
def route(query: str) -> str:
    complexity = estimate_complexity(query)   # rules / small classifier / heuristics
    model = "qwen2.5-7b-awq" if complexity < 0.5 else "qwen2.5-32b-awq"
    return call_llm(model, query)

5. Security and Compliance: Deployment's Last Gate ​

The moment your model goes live, it presents an attack surface. At minimum, do these four things (full governance framework in AI Safety and Governance):

MeasureHowWhere it applies
Content filteringAttach a moderation model/rules on the output side to block policy-violating contentConsumer-facing chat and writing products
Data maskingOn the request side, detect and replace PII first (phone numbers, ID numbers, emails, etc.), then restore it after the responseAny scenario handling users' private data
Audit loggingRecord inputs/outputs, model version, time, user identity; leave a trail for sensitive operationsCompliance audits, incident tracing
Permissions and isolationFor agent-style applications, restrict tool permissions and require human review of high-risk operationsAgent applications and tool-calling scenarios
python
# Minimal masking example: replace sensitive data before the request hits the model, restore it after the response
import re

def mask_pii(text: str):
    text = re.sub(r"1[3-9]\d{9}", "[phone number]", text)          # phone number
    text = re.sub(r"\d{17}[\dXx]", "[ID number]", text)        # ID card number
    return text

Self-hosting does not mean you're safe by default

Self-hosting only keeps data inside your intranet — content safety and prompt-injection defenses must not be skipped at all. Prompt injection, jailbreaks, and similar attacks work just as well against self-hosted models, and may even land more easily because the model is more "exposed". Security is a precondition for going live, not a patch applied afterward.

6. Common Pitfalls: Read This List Before You Ship ​

PitfallSymptomPrevention and fix
VRAM OOMCrash at startup or CUDA out of memory at runtimeQuantize before launch; start gpu-memory-utilization at 0.85 and tune up; budget VRAM for the KV cache and max-model-len together
TimeoutsThe client times out before the server does; users see "no response"Streaming output lowers perceived first-token latency; relax the gateway's proxy_read_timeout; size client timeouts from max_tokens
Concurrency exhaustedOne user's batch of requests occupies every concurrency slot while everyone else queuesGateway rate limiting + per-user quotas; cap it with max-num-seqs; queueing and backpressure as the safety net
Runaway output lengthAnswers balloon endlessly, token bills explode, APIs time outAdd length constraints to prompts (see Prompt Engineering Playbook); enforce a hard max_tokens cap; truncate overlong outputs
Hallucinations reaching productionAnswers confidently fabricate facts; user complaintsGround with RAG + "answer only from the provided material" (Building a RAG App from Scratch); add output-side validation
Silent accuracy regressionQuality degrades after quantization; users churnCompare on the same eval set before and after quantization (Setting Up LLM Evaluation)
No version managementAfter switching model/quantization versions, you can't roll backPut model files and configs under version control; alias with served-model-name for one-click switching

For a more complete list of "anti-patterns" (including evaluation leakage, data contamination, and other pitfalls beyond deployment), see Common Pitfalls and Anti-Patterns.

Further Reading ​

References ​