Skip to content

Inference Optimization and Quantization

At a glance Train once, infer countless times — inference is the biggest bill an LLM application will ever pay; this article breaks down the core inference optimization techniques, including KV Cache, INT4 quantization, distillation, speculative decoding, continuous batching, and FlashAttention, and closes with a serving framework comparison and an end-to-end tuning path.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Inference Optimization and Quantization ​

Inference optimization is the collection of techniques — quantization, caching, batching, distillation, and more — that reduce the latency and cost of large language model (LLM) inference while raising throughput. If training is "building the engine," inference optimization is "running that engine on the highway at minimum fuel consumption" — it determines whether your LLM application can scale profitably, whether it is fast enough, and whether it stays stable.

The one-line takeaway: inference is the biggest bill an LLM application pays; training costs money once, but every single response costs money. This article starts with the economics, then dissects six core techniques one by one — KV Cache, quantization, distillation, speculative decoding, batching, and structural optimization — and closes with a framework selection table and an end-to-end tuning path.

1. Why Inference Optimization Matters ​

1. The ledger: train once, infer countless times ​

Training and inference have completely different cost structures. Training is a one-time investment, while inference charges you on every call:

StageCost magnitude (approx., as of 2025-06)Frequency
PretrainingTens of millions of dollars in GPU compute for GPT-4-class modelsOne-time
Fine-tuning (PEFT)Hundreds to thousands of dollars (days on a single GPU)Occasional
InferenceRoughly $0.1–$0.8 per million input tokens and $0.4–$4 per million output tokens (depends on model and vendor)Every second

A chat application with 100,000 daily active users generates hundreds of millions of tokens a day — a single day of inference can cost more than one round of fine-tuning a small model. That is why the industry often says "training is expensive per run, inference is expensive in aggregate" — and the larger the large language model (LLM), the crueler this rule gets.

2. Latency decides the product experience ​

Inference optimization is not just about saving money — it directly decides whether a product lives or dies:

  • TTFT (time to first token): the time from the user pressing Enter to the first character appearing; beyond 2 seconds, churn rises noticeably;
  • TPOT (time per output token): the speed of the typewriter-style, token-by-token output, which determines how "smooth" it feels;
  • Throughput: how many requests a single card or cluster can serve concurrently, which determines unit cost.

Chat and search products like ChatGPT and Perplexity became the benchmarks they are thanks to sustained investment in inference optimization — see ChatGPT and Conversational AI. And DeepSeek-R1 went a step further, turning "low-cost, high-throughput inference" into a core competitive advantage.

2. Core Techniques at a Glance ​

Inference optimization falls broadly into two categories: the "algorithm layer" and the "systems layer." Here is the full picture first, then each technique in detail:

TechniqueCategoryOptimization goalTypical gain (approx.)
KV CacheAlgorithm/systemsAvoid recomputation, save VRAM10–100x faster inference (vs. recomputing at every step)
Quantization (INT8/INT4)AlgorithmShrink model size and computeVRAM cut by half or more; 1.5–3x faster
DistillationAlgorithmMatch a large model's quality with a small one5–50x smaller; much lower latency
Speculative decodingAlgorithmCut serial token-by-token latency1.5–3x end-to-end speedup
Continuous batchingSystemsRaise GPU utilization2–10x throughput
PagedAttentionSystemsEliminate KV cache fragmentationVRAM waste drops from 60–80% to <4%
FlashAttentionSystemsReduce VRAM reads/writes2–4x faster attention
Sparse attentionAlgorithmLower attention complexity for long inputsLarge drops in long-context VRAM and time

3. KV Cache: Making Repeated Computation Disappear ​

Transformer autoregressive decoding is inherently "generate one token at a time": every new token requires re-running attention over the entire history of tokens. The KV Cache is the caching technique that cuts through this bottleneck.

The prerequisite for understanding it is the attention mechanism: for every token, the attention layer computes three vectors — Query (Q), Key (K), and Value (V) — and each new token must dot against the K and V of all past tokens. Without a cache, step N must recompute the K and V of the previous N-1 tokens; with a KV cache, historical K and V are stored in VRAM, and each step only computes the new token's K and V and appends them to the cache — turning O(sequence length²) of repeated computation into O(1) appends.

No cache:    step N  compute K, V for tokens 1..N  → attention with Q           ← everything before was wasted
With cache:  step N  compute K, V for token N only → append to the KV Cache    → run attention directly

KV cache memory grows proportionally with context length and is the single biggest variable in inference memory. For a 7B model, a single request with a 2048-token context needs roughly 0.5GB of VRAM for the KV cache, and it grows linearly with batch size and sequence length — which leads directly to PagedAttention below (see Section 7). The mechanics of the KV cache rest entirely on attention; revisit Transformers and Attention for the fundamentals.

The one-line takeaway

The KV cache is the classic trade of "memory for speed" — it spares the generation phase from repeated computation, but memory becomes the new bottleneck. Every LLM serving framework enables it by default; all you need to worry about is how to store it and how to economize on it.

4. Quantization: Putting the Weights on a Diet ​

Quantization represents model weights (and sometimes activations as well) with fewer bits, reducing memory usage and speeding up computation. Every halving of weight precision roughly halves the model size:

PrecisionBytes per weight7B model weight sizevs. FP16Typical effect
FP324~28GB200%Training and high-precision baseline
FP16/BF162~14GB100%Default inference precision
INT81~7GB50%Negligible quality loss
INT40.5~3.5GB25%Slight loss; acceptable for most tasks

1. Weight quantization: GPTQ and AWQ ​

The most common approach in LLM inference is post-training quantization (PTQ) — no retraining required; you take the pretrained weights and represent them in low bits directly. The mainstream methods:

  • GPTQ (2022, arXiv:2210.17323): layer-wise quantization that uses second-order (Hessian) information for error compensation, applying a second-order approximation to correct quantization error within each layer. The perplexity loss after 4-bit quantization is small, making it the classic choice for INT4 deployment in the open-source community.
  • AWQ (Activation-aware Weight Quantization, 2023, arXiv:2306.00978): identifies "important channels" based on activation distributions, keeps them at higher precision, and quantizes the remaining channels to INT3/INT4. It preserves accuracy without relying on recalibration, which makes it deployment-friendly.

2. Dynamic vs. static quantization ​

DimensionStatic quantizationDynamic quantization
WeightsQuantized offlineQuantized offline
ActivationsFixed ranges computed from a calibration setRanges computed on the fly at runtime
SpeedFaster (known ranges can be precomputed)Slightly slower (extra runtime overhead)
AccuracyDepends on how representative the calibration set isMore robust (actual per-layer ranges)
Best forFixed cloud workloadsCPU inference; workloads with volatile input distributions

3. Quantization is not free ​

Quantization saves memory and speeds things up, but some accuracy is always sacrificed. How do you verify quality? You must run comparisons on LLM evaluation and benchmarks that matter to your business — not just perplexity:

Three pitfalls of quantization

  1. Perplexity alone will deceive you: a small perplexity loss does not mean math, code, or long-context tasks have not degraded — compare against an FP16 baseline on real task sets;
  2. Activation quantization is riskier than weight quantization: activations swing wildly with the input, and outlier tokens can be "clipped" into large errors;
  3. Not every layer should be quantized uniformly: the first layer, the last layer, and the attention projection layers are usually more sensitive; AWQ's approach (per-channel weighting) is more stable than one-size-fits-all.

One-line takeaway: start with INT8 (nearly lossless); move to INT4 only if you must (with evaluations as a safety net) — for other levers such as distillation, see the next section.

5. Distillation: A Large Model Teaches a Small One ​

Knowledge distillation has a "teacher" large model teach a "student" small model: during training, the student learns not only the teacher's correct answers (hard labels) but also the teacher's probability distributions (soft labels), thereby inheriting the generalizations the large model has learned. Hinton et al.'s classic 2015 paper framed it as "distilling knowledge from a large model into a small one."

Distillation vs. quantization ​

DimensionQuantizationDistillation
Model structureUnchanged; only numerical precision changesReplaced with a smaller architecture / fewer parameters
Memory gain4–8x (INT4)Tens of times (depends on the small model's scale)
Quality lossUsually smallStrongly correlated with the compression ratio
CostNearly free (PTQ requires no training)Requires training (can reuse data generated by the large model)
Relation to fine-tuningIndependentCan be combined with Fine-Tuning and PEFT (LoRA)

The two can be stacked: distill a small model first, then apply INT4 quantization to it — the standard playbook for mobile and on-device scenarios. A side effect of distillation is the "data flywheel": the teacher's outputs can be recycled as training data for iteration after iteration, getting stronger with use.

When to choose distillation

  • When you need extreme low latency / on-device deployment (say, running on a phone) → distill;
  • When you simply want to save cloud costs with minimal changes → quantize first;
  • Mixed budget → quantize first (live in a day), then distill once bottlenecks emerge (weeks).

6. Speculative Decoding: The Small Model Drafts, the Large Model Checks ​

Autoregressive generation is serial: every token must wait for the previous one to finish, so GPU utilization is naturally low. The idea behind speculative decoding is "let a fast small model draft, then let the large model verify in parallel":

  1. The draft model (a small model, usually around 1B scale) quickly generates n candidate tokens;
  2. The target model (the large model) runs one parallel forward pass to compute the true probabilities for those n candidates;
  3. Rejection sampling: accept correct tokens from the draft according to probability and roll back on any mismatch — since verifying n tokens in parallel takes only one forward pass, end-to-end latency drops by 1.5–3x, and the output distribution is exactly identical to running without speculation (its single biggest virtue).
Draft:   "The weather today is really" → generates ["nice", "!", "And"]
Verify:  the large model scores all three positions in one forward pass → all accepted → 3 tokens "earned" per forward pass

The correctness of speculative decoding is theoretically guaranteed (its output follows the same distribution as plain autoregressive sampling), so it costs virtually nothing in quality. It suits streaming chat, code completion, and other scenarios sensitive to per-token latency. Note that it optimizes latency, not throughput, and it requires loading an extra draft model, so memory usage increases slightly.

7. Batching: The Engine of Throughput ​

1. Continuous batching ​

Traditional serving frameworks schedule by "request batches": a batch either finishes together or waits together for the slowest one — slow requests drag down the entire batch (head-of-line blocking) and throughput suffers. Continuous batching (also called iteration-level scheduling) refines the scheduling granularity from "request" to "one forward iteration": at every step it dynamically decides which requests participate in compute and which can join or leave, keeping GPU memory and compute saturated. Compared with request-level batching, throughput gains of 2–10x are possible — one of the core reasons modern frameworks like vLLM crush older solutions on throughput.

2. PagedAttention: Paging the KV cache ​

The KV cache grows with requests; preallocating large blocks wastes memory, while dynamic allocation produces fragmentation. PagedAttention (the core innovation of vLLM, see paper arXiv:2309.06180) borrows the idea of virtual memory paging from operating systems:

  • It splits the KV cache into fixed-size "pages" (blocks) and uses a page table to map logically contiguous storage onto physically scattered memory blocks;
  • Pages are allocated on demand as requests grow, and neighboring requests can share pages with identical prefixes (such as the system prompt);
  • Memory waste drops from 60–80% under traditional schemes to under 4%, fitting more requests into the same batch.

This is also why "shared prefix" scenarios — RAG applications that inject the same document first, and the highly similar query prefixes of Retrieval-Augmented Generation (RAG) and Vector Databases and Semantic Search — see explosive throughput gains.

8. Structural Optimization: FlashAttention and Sparse Attention ​

1. FlashAttention: Keeping attention from round-tripping through VRAM ​

Standard attention has to write the entire N×N attention matrix to and read it back from VRAM (HBM), which makes memory bandwidth the bottleneck. FlashAttention (2022, arXiv:2205.14135) uses IO-aware tiling to perform the computation in blocks inside SRAM, avoiding materializing the full attention matrix: training gets 2–4x faster, and memory drops from O(N²) to O(N). FlashAttention-2 (2023) pushed parallelism and warp partitioning close to the theoretical peak — today every mainstream framework integrates it by default, making it an invisible optimization you "use without noticing."

2. Sparse attention ​

Attention complexity is O(N²), where N is the sequence length. Sparse attention makes each token attend only to "a local window plus a handful of global tokens," reducing complexity to roughly O(N). Representative approaches include the sliding windows plus global sentinel tokens of Longformer and BigBird. The cost is potential information loss and degraded long-context quality; much frontier work is exploring "sparse + dense hybrids" to get the best of both worlds — see Frontier Research.

ApproachComplexityQuality impactBest for
Full attentionO(N²)LosslessGeneral baseline
FlashAttentionO(N²) but bandwidth-optimizedLosslessAll scenarios (on by default)
Sparse attention~O(N)Lossy (long-range dependencies weakened)Very long contexts, edge deployment

9. Choosing a Serving Framework ​

The same optimization techniques vary enormously in implementation depth from one framework to another. The mainstream choices as of 2025-06:

FrameworkKey strengthsPositioningLearning curveTypical use
vLLMPagedAttention, continuous batching, OpenAI API compatibleThe main open-source workhorse for the cloudMediumProduction API services, RAG/agent backends
TensorRT-LLMFrom NVIDIA; fused TensorRT kernels, deeply tied to GPUsPeak cloud performanceHighNVIDIA GPU production environments, maximum throughput
llama.cppGGUF quantization format, CPU/Apple Silicon optimizedLocal/edgeLowGPU-less machines, on-device, personal deployment
SGLangRadixAttention prefix reuse, structured output, agent scenariosEmerging cloud optionMediumMulti-turn chat/agents, constrained JSON output
OllamaOne-click experience built on llama.cppLocal personal useLowestLocal trials, developer toys, education

One-line guidance: for a public-facing service pick vLLM (the most mature ecosystem) or TensorRT-LLM (squeezes NVIDIA hardware dry); for your own machine pick llama.cpp/Ollama; if agents and structured generation dominate your workload, evaluate SGLang. For deployment details, see LLM Deployment and Inference Optimization in Practice.

10. The End-to-End Tuning Path ​

Optimization is not about piling on techniques; it is about making decisions in order of "bang for the buck." The recommended end-to-end path:

StepActionGainPitfall
1. Model choiceStart with the smallest model that meets the quality bar (see Model and Leaderboard Quick Reference)Save money at the sourceJumping straight to the biggest model
2. QuantizationApply INT8/INT4 (GPTQ/AWQ)VRAM halved or betterShipping without evaluation
3. CachingEnable the KV cache, shared prefixes/prompt cachingBig latency dropIgnoring long-context memory footprint
4. BatchingContinuous batching + PagedAttention-class frameworks2–10x throughputGuessing batch sizes
5. HardwarePick GPUs by memory and bandwidth (bandwidth determines decode speed)Higher per-card throughputLooking at memory but not bandwidth
6. DeploymentA vLLM-class framework + monitoring latency/throughput metricsOperable and scalableNo monitoring after launch

A rule of thumb for hardware: the decode (generation) phase is memory-bandwidth-bound, so bandwidth-oriented GPUs (H200, A100, etc.) suit generation-heavy workloads better than pure compute ones; compute only becomes the bottleneck under high request concurrency. For the complete battle manual for your team, see LLM Deployment and Inference Optimization in Practice.

Baseline first, optimize second

Every optimization starts with a baseline: on the same dataset, measure the unoptimized TTFT, TPOT, tokens/s, and per-token cost; then add optimizations one at a time and compare each against the last. Do not turn everything on at once — every optimization has diminishing marginal returns; fix the biggest bottleneck first.

11. Trade-offs ​

1. The speed vs. quality vs. cost triangle ​

Trade-offNotes
Speed ↔ QualityINT4, distillation, and speculative drafts can all degrade or alter output quality; use evaluations as the safety net
Speed ↔ CostHigh-throughput frameworks and clusters need total-cost accounting: lower per-token cost, higher engineering and hardware spend
Quality ↔ CostLarge models are high quality but expensive; distillation is precisely trading training cost for inference cost
Latency ↔ ThroughputSpeculative decoding cuts latency but adds memory; continuous batching raises throughput but slightly increases per-request latency

2. Let metrics speak, not gut feeling ​

The three yardsticks for measuring inference optimization (standard definitions in LLM Evaluation and Benchmarks; putting them into practice in Building an LLM Eval System):

MetricFull name / meaningWhat it measuresTypical target (chat scenario, approx.)
TTFTTime To First Token; first-token latencyHow quickly a response "gets off the line"< 1–2 seconds
TPOTTime Per Output Token; time per output tokenThe "cadence" of generation20–60ms/token
Throughputtokens/s (tokens generated per second)Output per unit timeHigher is better (lower cost)

Launch discipline after optimization

You must run regression evaluations before launch: quantization, distillation, and speculative decoding all change the output distribution, so run a round of business evaluations before releasing — this is especially critical for quantized models. For how to build an evaluation system, see Building an LLM Eval System; for interview and job-hunting perspectives, see Interview Questions. For consistent terminology, see Glossary.

One-line summary: the essence of inference optimization is delivering "model capability" at the lowest cost and lowest latency through engineering — quantize first, then batch, distill when necessary, and always verify with metrics.

Further Reading ​

References ​