Skip to content

LLM Inference Optimization

At a glance LLM inference is the hottest frontier in model deployment: autoregressive generation, KV cache, continuous batching, speculative decoding, tensor parallelism. This article lays out the GPU memory ledger of LLM serving and every optimization technique that matters.

LLM Inference Optimization ​

One-line definition: LLM inference optimization is the systems engineering of combining memory management, batch scheduling, parallelism, and quantization around one defining property — autoregressive, token-by-token generation — the goal is to make every generated token fast, cheap, and able to serve more concurrent requests.

Industry insight: LLM inference rewrote the playbook of traditional inference optimization. Your intuition that "the weights dominate GPU memory" is wrong in the LLM world — as concurrency and context length grow, the KV cache can eat more memory than the weights do. "Bigger batch means better economics" breaks down too: requests have different lengths, and static batching forces the whole batch to wait for the slowest one. vLLM made its name by turning "memory management," a long-ignored problem, into its core innovation (PagedAttention), proving that systems-level engineering can deliver order-of-magnitude gains. This section gives you a complete "GPU memory ledger + optimization toolbox."

1. Three Properties of LLM Inference ​

  1. Autoregressive generation: one forward pass produces one token, and each next token depends on the previous one — generating N tokens takes N forward passes; you can't compute it all at once;
  2. The KV cache eats memory: every request keeps caching attention keys and values as it generates, growing linearly with sequence length;
  3. Decoding is bandwidth-bound: each step computes only one new token, so shipping weights through memory dominates — which is exactly why quantization and batching are so effective for LLMs.
text
LLM generation timeline:
request: "deploy the model" + generating...
  Prefill: one forward pass over the full prompt → produces the first token (TTFT)
  Decode : one forward pass per step → one token (TPOT, repeated for every token)

2. The GPU Memory Ledger: Weights + Activations + KV Cache ​

Take a 7B FP16 model (weights ≈ 14GB) deployed on an 80GB A100. The memory budget breaks down like this:

text
┌────────────────────────────────────────────────────────────────┐
│  Weights ~14GB (FP16)                                          │
│  Activations ~2GB (at batch=16)                                │
│  KV cache: grows with concurrency × context length             │
│    per request, per token ≈ 2 × layers × heads × head_dim × 2B │
│    example: 32 layers × 128 heads × 128 dim × 2B ≈ 2MB/token   │
│    128 concurrent requests, 2K tokens context on average       │
│    ≈ 128 × 2048 × 2MB = 512GB (far beyond the weights!)        │
└────────────────────────────────────────────────────────────────┘

The key takeaway: the KV cache is not a minor overhead — it is the dominant memory item. A formula for estimating KV cache size:

text
KV cache size = 2 × layers × num_kv_heads × head_dim × bytes (FP16=2) × total sequence length

That's why "how long a context and how many concurrent requests can we support" is the central memory decision in LLM deployment. Mitigations: KV cache quantization (INT8/FP4), GQA/MQA architectures (shared KV heads), and the PagedAttention-style dynamic management covered next.

3. Seven Core Optimizations ​

1. KV Cache Management: PagedAttention ​

The core idea of the vLLM paper (SOSP 2023): manage the KV cache the way an operating system manages memory — allocate fixed-size blocks (pages), link them through an index table, and eliminate the fragmentation waste of "reserving a large contiguous region of GPU memory but using only a fraction of it." The payoffs:

  • Memory utilization climbs from ~60% to 90%+;
  • The same page can be shared across requests (e.g., common prefixes in beam search);
  • Throughput gains of up to 20×+ (measured in the paper, long-context scenarios). For a deep dive into the paper, see Papers: PagedAttention.

2. Continuous Batching ​

With static batching, a batch holds its resources until every request in it finishes, so short requests wait on long ones. Continuous batching switches to token-level scheduling:

text
Static batching:      request A ████████
                      request B ████████████  ← B is short but must wait for A; the GPU idles
Continuous batching:  request A ████ ✓ (finishes first, leaves first)
                      request B ████████
                      request C       ████████  ← new requests slot into freed capacity immediately

This is the number-one lever for multiplying LLM serving throughput several times over, and vLLM, TGI, and SGLang all ship it built in.

3. Speculative Decoding ​

The idea: a small draft model guesses the next K tokens in one shot, then the large model verifies all of them in a single forward pass — correct guesses are pure profit; wrong ones get rolled back. When QPS isn't your constraint (you have spare compute), speedups of 2–3× are achievable, and the output matches greedy decoding exactly.

text
Draft model: guesses 4 tokens, e.g. "hello there world today"
Large model, one forward pass: verifies all 4 positions at once
  → 3 correct, 1 wrong → accept the first 3, redo from the mistake
  → on average each forward pass "earns" 2–3 tokens for free

4. Quantization: The Bandwidth Fix ​

Every decode step streams the full set of weights through memory, so weight precision directly determines speed: FP16 → INT4 cuts weight traffic by 75%, with near-linear gains in generation speed. For quantization methods and calibration discipline, see Quantization; the key methods are AWQ/GPTQ/GGUF.

5. Prefix Caching (Prompt Caching) ​

LLM prompts often carry long, unchanging blocks of system instructions and context. Cache the KV results of the shared prefix, and every request that hits the cache skips prefill — TTFT can drop by 50% or more. Multi-turn conversations benefit the most. vLLM's prefix caching and OpenAI's prompt caching are both this idea.

6. Prefill/Decode Disaggregation (PD) ​

Prefill (compute-bound — it craves FLOPs) and decode (memory-bound — it craves bandwidth) stress hardware differently. Dedicate one pool of GPUs to prefill and another to decode, with scheduling in between, and both resource types can run at full utilization. This suits very large scale and high concurrency; it's an advanced technique.

7. Memory Offloading ​

Weights and activations move between GPU and CPU memory on demand, trading roughly an order of magnitude in speed for the ability to run a large model on a small GPU — only suitable for "it just needs to run" scenarios (e.g., local tools).

4. Parallelism: Tensor / Pipeline / Data ​

SchemeWhat gets splitCommunication frequencyBest for
Tensor parallel (TP)Each layer's matrices across GPUsMultiple times per layer (very high)Model too big for one GPU, 2–8 GPUs
Pipeline parallel (PP)Layers grouped onto GPUsOnce per stage (low)Very deep models, cross-node
Data parallel (DP)Full model × N GPUsOnly at synchronizationThroughput scaling, combined with TP/PP

Engineering note: TP is communication-hungry and needs NVLink/InfiniBand-class interconnects; across nodes, prefer PP over TP. For a deeper treatment, see Papers: Parallel and Distributed Inference.

text
Typical 70B deployment on 8×H100:
  TP=8 (tensor split within one node, fully connected over NVLink)
  or TP=4 × PP=2 (2 nodes, low inter-node communication)
  Let load tests of your actual communication bandwidth make the final call

5. Serving Frameworks at a Glance ​

FrameworkPositioningHighlights
vLLMHigh-throughput LLM servingPagedAttention + continuous batching + prefix caching
Hugging Face TGIDeep HF ecosystem integrationEasy to deploy, feature-complete
SGLangExtreme scheduling (RadixAttention)Strongest performance on complex prompt workloads
llama.cppLocal/CPU friendlyCross-platform, GGUF ecosystem
TensorRT-LLMThe limit of NVIDIA hardwareGraph compilation + a full set of integrated optimizations

How to choose: balance ecosystem against performance. vLLM is currently the default choice with the widest production adoption; for hands-on experience, see LLM Serving with vLLM.

6. Performance Metrics and Cost Optimization ​

  • TTFT: set by prefill speed; affected by input length and prefill compute;
  • TPOT: set by decode speed; affected by weight precision and memory bandwidth;
  • Throughput (tokens/s): the direct beneficiary of continuous batching + quantization.

Cost optimization priority order:

text
① Continuous batching (2–5× throughput, nearly free)
② Quantize to INT8/INT4 (halve weight traffic / cut it by 75%)
③ KV cache quantization + GQA (more concurrency / longer context)
④ Prefix caching (big TTFT drop in conversational workloads)
⑤ Speculative decoding (when you have QPS headroom)
⑥ Elastic scaling + PD disaggregation (at larger scale)

7. A Typical Deployment Architecture ​

text
Clients
  │ streaming SSE/WebSocket
  ▼
API gateway (auth / rate limiting / routing)
  ▼
LLM inference cluster (K8s)
  ├─ N × vLLM instances (each can load multiple models, routed by model)
  │    └─ tensor parallelism (multi-GPU)
  ▼
Monitoring: TTFT/TPOT/throughput/GPU/memory (see /concepts/monitoring)

Trade-offs ​

DecisionOptionsHow to choose
Throughput vs. latencyLarge batch vs. small batchProtect TTFT for interactive traffic; chase throughput for offline generation
FrameworkvLLM vs. TensorRT-LLM vs. llama.cppvLLM for production at high concurrency; TensorRT-LLM to max out NVIDIA hardware; llama.cpp for local
Quantization precisionFP16 vs. INT8 vs. INT4Drop precision only when bandwidth-bound, and always validate quality
Context lengthLonger means more KV cacheCap it at what the business actually needs; don't chase 128K blindly
Single GPU vs. multi-GPUSimplicity vs. capacity/throughputQuantize first → then TP → then PP

In one sentence: LLM inference engineering boils down to balancing three ledgers — the memory ledger (the KV cache is the star), the scheduling ledger (continuous batching keeps every moment busy), and the communication ledger (don't let interconnects strangle your parallelism). Balance all three, and throughput and cost fall into place.

Further Reading ​

References ​