Appearance
Inference Optimization and Quantization
Inference optimization is the collection of techniques — quantization, caching, batching, distillation, and more — that reduce the latency and cost of large language model (LLM) inference while raising throughput. If training is "building the engine," inference optimization is "running that engine on the highway at minimum fuel consumption" — it determines whether your LLM application can scale profitably, whether it is fast enough, and whether it stays stable.
The one-line takeaway: inference is the biggest bill an LLM application pays; training costs money once, but every single response costs money. This article starts with the economics, then dissects six core techniques one by one — KV Cache, quantization, distillation, speculative decoding, batching, and structural optimization — and closes with a framework selection table and an end-to-end tuning path.
1. Why Inference Optimization Matters
1. The ledger: train once, infer countless times
Training and inference have completely different cost structures. Training is a one-time investment, while inference charges you on every call:
| Stage | Cost magnitude (approx., as of 2025-06) | Frequency |
|---|---|---|
| Pretraining | Tens of millions of dollars in GPU compute for GPT-4-class models | One-time |
| Fine-tuning (PEFT) | Hundreds to thousands of dollars (days on a single GPU) | Occasional |
| Inference | Roughly $0.1–$0.8 per million input tokens and $0.4–$4 per million output tokens (depends on model and vendor) | Every second |
A chat application with 100,000 daily active users generates hundreds of millions of tokens a day — a single day of inference can cost more than one round of fine-tuning a small model. That is why the industry often says "training is expensive per run, inference is expensive in aggregate" — and the larger the large language model (LLM), the crueler this rule gets.
2. Latency decides the product experience
Inference optimization is not just about saving money — it directly decides whether a product lives or dies:
- TTFT (time to first token): the time from the user pressing Enter to the first character appearing; beyond 2 seconds, churn rises noticeably;
- TPOT (time per output token): the speed of the typewriter-style, token-by-token output, which determines how "smooth" it feels;
- Throughput: how many requests a single card or cluster can serve concurrently, which determines unit cost.
Chat and search products like ChatGPT and Perplexity became the benchmarks they are thanks to sustained investment in inference optimization — see ChatGPT and Conversational AI. And DeepSeek-R1 went a step further, turning "low-cost, high-throughput inference" into a core competitive advantage.
2. Core Techniques at a Glance
Inference optimization falls broadly into two categories: the "algorithm layer" and the "systems layer." Here is the full picture first, then each technique in detail:
| Technique | Category | Optimization goal | Typical gain (approx.) |
|---|---|---|---|
| KV Cache | Algorithm/systems | Avoid recomputation, save VRAM | 10–100x faster inference (vs. recomputing at every step) |
| Quantization (INT8/INT4) | Algorithm | Shrink model size and compute | VRAM cut by half or more; 1.5–3x faster |
| Distillation | Algorithm | Match a large model's quality with a small one | 5–50x smaller; much lower latency |
| Speculative decoding | Algorithm | Cut serial token-by-token latency | 1.5–3x end-to-end speedup |
| Continuous batching | Systems | Raise GPU utilization | 2–10x throughput |
| PagedAttention | Systems | Eliminate KV cache fragmentation | VRAM waste drops from 60–80% to <4% |
| FlashAttention | Systems | Reduce VRAM reads/writes | 2–4x faster attention |
| Sparse attention | Algorithm | Lower attention complexity for long inputs | Large drops in long-context VRAM and time |
3. KV Cache: Making Repeated Computation Disappear
Transformer autoregressive decoding is inherently "generate one token at a time": every new token requires re-running attention over the entire history of tokens. The KV Cache is the caching technique that cuts through this bottleneck.
The prerequisite for understanding it is the attention mechanism: for every token, the attention layer computes three vectors — Query (Q), Key (K), and Value (V) — and each new token must dot against the K and V of all past tokens. Without a cache, step N must recompute the K and V of the previous N-1 tokens; with a KV cache, historical K and V are stored in VRAM, and each step only computes the new token's K and V and appends them to the cache — turning O(sequence length²) of repeated computation into O(1) appends.
No cache: step N compute K, V for tokens 1..N → attention with Q ← everything before was wasted
With cache: step N compute K, V for token N only → append to the KV Cache → run attention directlyKV cache memory grows proportionally with context length and is the single biggest variable in inference memory. For a 7B model, a single request with a 2048-token context needs roughly 0.5GB of VRAM for the KV cache, and it grows linearly with batch size and sequence length — which leads directly to PagedAttention below (see Section 7). The mechanics of the KV cache rest entirely on attention; revisit Transformers and Attention for the fundamentals.
The one-line takeaway
The KV cache is the classic trade of "memory for speed" — it spares the generation phase from repeated computation, but memory becomes the new bottleneck. Every LLM serving framework enables it by default; all you need to worry about is how to store it and how to economize on it.
4. Quantization: Putting the Weights on a Diet
Quantization represents model weights (and sometimes activations as well) with fewer bits, reducing memory usage and speeding up computation. Every halving of weight precision roughly halves the model size:
| Precision | Bytes per weight | 7B model weight size | vs. FP16 | Typical effect |
|---|---|---|---|---|
| FP32 | 4 | ~28GB | 200% | Training and high-precision baseline |
| FP16/BF16 | 2 | ~14GB | 100% | Default inference precision |
| INT8 | 1 | ~7GB | 50% | Negligible quality loss |
| INT4 | 0.5 | ~3.5GB | 25% | Slight loss; acceptable for most tasks |
1. Weight quantization: GPTQ and AWQ
The most common approach in LLM inference is post-training quantization (PTQ) — no retraining required; you take the pretrained weights and represent them in low bits directly. The mainstream methods:
- GPTQ (2022, arXiv:2210.17323): layer-wise quantization that uses second-order (Hessian) information for error compensation, applying a second-order approximation to correct quantization error within each layer. The perplexity loss after 4-bit quantization is small, making it the classic choice for INT4 deployment in the open-source community.
- AWQ (Activation-aware Weight Quantization, 2023, arXiv:2306.00978): identifies "important channels" based on activation distributions, keeps them at higher precision, and quantizes the remaining channels to INT3/INT4. It preserves accuracy without relying on recalibration, which makes it deployment-friendly.
2. Dynamic vs. static quantization
| Dimension | Static quantization | Dynamic quantization |
|---|---|---|
| Weights | Quantized offline | Quantized offline |
| Activations | Fixed ranges computed from a calibration set | Ranges computed on the fly at runtime |
| Speed | Faster (known ranges can be precomputed) | Slightly slower (extra runtime overhead) |
| Accuracy | Depends on how representative the calibration set is | More robust (actual per-layer ranges) |
| Best for | Fixed cloud workloads | CPU inference; workloads with volatile input distributions |
3. Quantization is not free
Quantization saves memory and speeds things up, but some accuracy is always sacrificed. How do you verify quality? You must run comparisons on LLM evaluation and benchmarks that matter to your business — not just perplexity:
Three pitfalls of quantization
- Perplexity alone will deceive you: a small perplexity loss does not mean math, code, or long-context tasks have not degraded — compare against an FP16 baseline on real task sets;
- Activation quantization is riskier than weight quantization: activations swing wildly with the input, and outlier tokens can be "clipped" into large errors;
- Not every layer should be quantized uniformly: the first layer, the last layer, and the attention projection layers are usually more sensitive; AWQ's approach (per-channel weighting) is more stable than one-size-fits-all.
One-line takeaway: start with INT8 (nearly lossless); move to INT4 only if you must (with evaluations as a safety net) — for other levers such as distillation, see the next section.
5. Distillation: A Large Model Teaches a Small One
Knowledge distillation has a "teacher" large model teach a "student" small model: during training, the student learns not only the teacher's correct answers (hard labels) but also the teacher's probability distributions (soft labels), thereby inheriting the generalizations the large model has learned. Hinton et al.'s classic 2015 paper framed it as "distilling knowledge from a large model into a small one."
Distillation vs. quantization
| Dimension | Quantization | Distillation |
|---|---|---|
| Model structure | Unchanged; only numerical precision changes | Replaced with a smaller architecture / fewer parameters |
| Memory gain | 4–8x (INT4) | Tens of times (depends on the small model's scale) |
| Quality loss | Usually small | Strongly correlated with the compression ratio |
| Cost | Nearly free (PTQ requires no training) | Requires training (can reuse data generated by the large model) |
| Relation to fine-tuning | Independent | Can be combined with Fine-Tuning and PEFT (LoRA) |
The two can be stacked: distill a small model first, then apply INT4 quantization to it — the standard playbook for mobile and on-device scenarios. A side effect of distillation is the "data flywheel": the teacher's outputs can be recycled as training data for iteration after iteration, getting stronger with use.
When to choose distillation
- When you need extreme low latency / on-device deployment (say, running on a phone) → distill;
- When you simply want to save cloud costs with minimal changes → quantize first;
- Mixed budget → quantize first (live in a day), then distill once bottlenecks emerge (weeks).
6. Speculative Decoding: The Small Model Drafts, the Large Model Checks
Autoregressive generation is serial: every token must wait for the previous one to finish, so GPU utilization is naturally low. The idea behind speculative decoding is "let a fast small model draft, then let the large model verify in parallel":
- The draft model (a small model, usually around 1B scale) quickly generates n candidate tokens;
- The target model (the large model) runs one parallel forward pass to compute the true probabilities for those n candidates;
- Rejection sampling: accept correct tokens from the draft according to probability and roll back on any mismatch — since verifying n tokens in parallel takes only one forward pass, end-to-end latency drops by 1.5–3x, and the output distribution is exactly identical to running without speculation (its single biggest virtue).
Draft: "The weather today is really" → generates ["nice", "!", "And"]
Verify: the large model scores all three positions in one forward pass → all accepted → 3 tokens "earned" per forward passThe correctness of speculative decoding is theoretically guaranteed (its output follows the same distribution as plain autoregressive sampling), so it costs virtually nothing in quality. It suits streaming chat, code completion, and other scenarios sensitive to per-token latency. Note that it optimizes latency, not throughput, and it requires loading an extra draft model, so memory usage increases slightly.
7. Batching: The Engine of Throughput
1. Continuous batching
Traditional serving frameworks schedule by "request batches": a batch either finishes together or waits together for the slowest one — slow requests drag down the entire batch (head-of-line blocking) and throughput suffers. Continuous batching (also called iteration-level scheduling) refines the scheduling granularity from "request" to "one forward iteration": at every step it dynamically decides which requests participate in compute and which can join or leave, keeping GPU memory and compute saturated. Compared with request-level batching, throughput gains of 2–10x are possible — one of the core reasons modern frameworks like vLLM crush older solutions on throughput.
2. PagedAttention: Paging the KV cache
The KV cache grows with requests; preallocating large blocks wastes memory, while dynamic allocation produces fragmentation. PagedAttention (the core innovation of vLLM, see paper arXiv:2309.06180) borrows the idea of virtual memory paging from operating systems:
- It splits the KV cache into fixed-size "pages" (blocks) and uses a page table to map logically contiguous storage onto physically scattered memory blocks;
- Pages are allocated on demand as requests grow, and neighboring requests can share pages with identical prefixes (such as the system prompt);
- Memory waste drops from 60–80% under traditional schemes to under 4%, fitting more requests into the same batch.
This is also why "shared prefix" scenarios — RAG applications that inject the same document first, and the highly similar query prefixes of Retrieval-Augmented Generation (RAG) and Vector Databases and Semantic Search — see explosive throughput gains.
8. Structural Optimization: FlashAttention and Sparse Attention
1. FlashAttention: Keeping attention from round-tripping through VRAM
Standard attention has to write the entire N×N attention matrix to and read it back from VRAM (HBM), which makes memory bandwidth the bottleneck. FlashAttention (2022, arXiv:2205.14135) uses IO-aware tiling to perform the computation in blocks inside SRAM, avoiding materializing the full attention matrix: training gets 2–4x faster, and memory drops from O(N²) to O(N). FlashAttention-2 (2023) pushed parallelism and warp partitioning close to the theoretical peak — today every mainstream framework integrates it by default, making it an invisible optimization you "use without noticing."
2. Sparse attention
Attention complexity is O(N²), where N is the sequence length. Sparse attention makes each token attend only to "a local window plus a handful of global tokens," reducing complexity to roughly O(N). Representative approaches include the sliding windows plus global sentinel tokens of Longformer and BigBird. The cost is potential information loss and degraded long-context quality; much frontier work is exploring "sparse + dense hybrids" to get the best of both worlds — see Frontier Research.
| Approach | Complexity | Quality impact | Best for |
|---|---|---|---|
| Full attention | O(N²) | Lossless | General baseline |
| FlashAttention | O(N²) but bandwidth-optimized | Lossless | All scenarios (on by default) |
| Sparse attention | ~O(N) | Lossy (long-range dependencies weakened) | Very long contexts, edge deployment |
9. Choosing a Serving Framework
The same optimization techniques vary enormously in implementation depth from one framework to another. The mainstream choices as of 2025-06:
| Framework | Key strengths | Positioning | Learning curve | Typical use |
|---|---|---|---|---|
| vLLM | PagedAttention, continuous batching, OpenAI API compatible | The main open-source workhorse for the cloud | Medium | Production API services, RAG/agent backends |
| TensorRT-LLM | From NVIDIA; fused TensorRT kernels, deeply tied to GPUs | Peak cloud performance | High | NVIDIA GPU production environments, maximum throughput |
| llama.cpp | GGUF quantization format, CPU/Apple Silicon optimized | Local/edge | Low | GPU-less machines, on-device, personal deployment |
| SGLang | RadixAttention prefix reuse, structured output, agent scenarios | Emerging cloud option | Medium | Multi-turn chat/agents, constrained JSON output |
| Ollama | One-click experience built on llama.cpp | Local personal use | Lowest | Local trials, developer toys, education |
One-line guidance: for a public-facing service pick vLLM (the most mature ecosystem) or TensorRT-LLM (squeezes NVIDIA hardware dry); for your own machine pick llama.cpp/Ollama; if agents and structured generation dominate your workload, evaluate SGLang. For deployment details, see LLM Deployment and Inference Optimization in Practice.
10. The End-to-End Tuning Path
Optimization is not about piling on techniques; it is about making decisions in order of "bang for the buck." The recommended end-to-end path:
| Step | Action | Gain | Pitfall |
|---|---|---|---|
| 1. Model choice | Start with the smallest model that meets the quality bar (see Model and Leaderboard Quick Reference) | Save money at the source | Jumping straight to the biggest model |
| 2. Quantization | Apply INT8/INT4 (GPTQ/AWQ) | VRAM halved or better | Shipping without evaluation |
| 3. Caching | Enable the KV cache, shared prefixes/prompt caching | Big latency drop | Ignoring long-context memory footprint |
| 4. Batching | Continuous batching + PagedAttention-class frameworks | 2–10x throughput | Guessing batch sizes |
| 5. Hardware | Pick GPUs by memory and bandwidth (bandwidth determines decode speed) | Higher per-card throughput | Looking at memory but not bandwidth |
| 6. Deployment | A vLLM-class framework + monitoring latency/throughput metrics | Operable and scalable | No monitoring after launch |
A rule of thumb for hardware: the decode (generation) phase is memory-bandwidth-bound, so bandwidth-oriented GPUs (H200, A100, etc.) suit generation-heavy workloads better than pure compute ones; compute only becomes the bottleneck under high request concurrency. For the complete battle manual for your team, see LLM Deployment and Inference Optimization in Practice.
Baseline first, optimize second
Every optimization starts with a baseline: on the same dataset, measure the unoptimized TTFT, TPOT, tokens/s, and per-token cost; then add optimizations one at a time and compare each against the last. Do not turn everything on at once — every optimization has diminishing marginal returns; fix the biggest bottleneck first.
11. Trade-offs
1. The speed vs. quality vs. cost triangle
| Trade-off | Notes |
|---|---|
| Speed ↔ Quality | INT4, distillation, and speculative drafts can all degrade or alter output quality; use evaluations as the safety net |
| Speed ↔ Cost | High-throughput frameworks and clusters need total-cost accounting: lower per-token cost, higher engineering and hardware spend |
| Quality ↔ Cost | Large models are high quality but expensive; distillation is precisely trading training cost for inference cost |
| Latency ↔ Throughput | Speculative decoding cuts latency but adds memory; continuous batching raises throughput but slightly increases per-request latency |
2. Let metrics speak, not gut feeling
The three yardsticks for measuring inference optimization (standard definitions in LLM Evaluation and Benchmarks; putting them into practice in Building an LLM Eval System):
| Metric | Full name / meaning | What it measures | Typical target (chat scenario, approx.) |
|---|---|---|---|
| TTFT | Time To First Token; first-token latency | How quickly a response "gets off the line" | < 1–2 seconds |
| TPOT | Time Per Output Token; time per output token | The "cadence" of generation | 20–60ms/token |
| Throughput | tokens/s (tokens generated per second) | Output per unit time | Higher is better (lower cost) |
Launch discipline after optimization
You must run regression evaluations before launch: quantization, distillation, and speculative decoding all change the output distribution, so run a round of business evaluations before releasing — this is especially critical for quantized models. For how to build an evaluation system, see Building an LLM Eval System; for interview and job-hunting perspectives, see Interview Questions. For consistent terminology, see Glossary.
One-line summary: the essence of inference optimization is delivering "model capability" at the lowest cost and lowest latency through engineering — quantize first, then batch, distill when necessary, and always verify with metrics.
Further Reading
- Transformers and Attention — the theoretical foundation of the KV cache and attention optimization
- Large Language Models (LLM) — what inference optimization ultimately serves
- LLM Evaluation and Benchmarks — how to verify quality after quantization/distillation
- Fine-Tuning and PEFT (LoRA) — the relationship between distillation and fine-tuning
- LLM Deployment and Inference Optimization in Practice — the complete manual from this article's theory to production
- Building an LLM Eval System — regression evaluation practice before launch
- Frontier Research — the latest work on FlashAttention, sparse attention, and other structural optimizations
- ChatGPT and Conversational AI — a product case study in large-scale inference serving
References
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention (2023, arXiv:2309.06180) — the original vLLM/PagedAttention paper
- Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (2022, arXiv:2210.17323) — the classic GPTQ weight quantization method
- Lin et al., AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (2023, arXiv:2306.00978) — AWQ activation-aware quantization
- Dao et al., FlashAttention: Fast and Efficient Exact Attention with IO-Awareness (2022, arXiv:2205.14135) — the original FlashAttention paper
- Leviathan et al., Fast Inference from Transformers via Speculative Decoding (2023, arXiv:2211.17192) — the theoretical foundation of speculative decoding
- vLLM documentation — docs for the production-grade serving framework
- llama.cpp repository — open-source implementation of GGUF quantization and local inference
- Ollama website — one-click local deployment tool
- NVIDIA TensorRT-LLM documentation — NVIDIA's peak-performance GPU inference stack