Theme
Inference Fundamentals: Autoregression and Sampling
Inference is the stage where the model truly "speaks" after training: given input, generating output token by token. Training determines "what the model knows"; inference configuration determines "how it says it" — the same model, temperature at 0 is a rigorous engineer, at 1.0 is a creative writer loose its constraints. That's the subject of this article.
This article unfolds along the path of "generation loop → KV Cache → decoding strategies → dialogue format → streaming output." Performance engineering for inference (quantization, batching, memory) is in Deployment & Serving; measuring sampling effects is in Evaluation & Benchmarks.
1. Autoregression: One Token at a Time
Most modern LLMs (GPT, Llama, Qwen, etc.) are autoregressive decoders: predict one next token at a time, append the new token to the input tail, then predict the next, looping until generation ends. This is the direct manifestation of language modeling's "predict the next word" paradigm at inference time.
text
Autoregressive generation loop (pseudocode):
Input tokens = [t1, t2, ..., tn]
Output tokens = []
while True:
# 1. Forward pass: get the next-token probability distribution
probs = model(tokens) # shape [vocab_size]
# 2. Pick a token per decoding strategy (greedy/sampling, see Section 3)
next_token = decode(probs)
# 3. Append to sequence
tokens.append(next_token)
outputs.append(next_token)
# 4. Stop at EOS or max_tokens
if next_token == EOS or len(outputs) >= max_tokens:
breakKey implication: generating n tokens requires n forward passes; step t must "see" all preceding t-1 tokens. This brings two engineering consequences — KV Cache (save repeated computation) and streaming output (send as you compute, no need to wait for full generation).
Why autoregression is the de facto standard
Autoregression reduces "generating text" to "repeatedly doing what language models do best (predict the next word)," aligning the training objective perfectly with the generation objective (by contrast, BERT-style masked models need extra design for generation, see Transformer Architecture Deep Dive). The cost is generation can't be parallelized — this is the root problem KV Cache and batching optimizations address.
2. KV Cache: Prefill and Decode Two Phases
In the attention mechanism, every token needs to compute Query, Key, Value (QKV). At step t, the new token must compute attention with all preceding tokens' Key/Value. Without caching, step t would need to recompute K/V for the preceding t-1 tokens — redundant work.
KV Cache caches the Key and Value of all preceding tokens in GPU memory, computing only the new token's K/V during generation, directly reusing the cache:
text
Prefill phase (prompt phase):
Process the full input prompt all at once, compute and cache K/V for all prompt tokens
── High parallelism, large throughput
Decode phase:
Generate token by token, each step only needs:
① Compute the new token's Q/K/V
② Compute attention between the new token's K/V and cached K/V
③ Append new K/V to cache
── Serial, latency-dominated (delay determined by single-step time)KV Cache is one of the main sources of memory overhead at inference time (in long context, it can even exceed weight memory), and is a key object in Context & Long Context and deployment optimization. Memory estimation, PagedAttention, etc. optimizations are in Deployment & Serving.
Long context ≠ unlimited KV Cache
KV Cache size grows linearly with sequence length; at long context (100K token range), memory pressure spikes dramatically. This is the gap between "model supports 128K context" and "can really run full 128K" — model capability, memory budget, and deployment config jointly determine the actual usable length.
The prefill and decode two phases have completely different meanings for user experience: prefill determines TTFT (first-token latency) — processing the full prompt all at once, short prompt fast, long prompt slow; decode determines TPOT (subsequent token interval) — token-by-token serial, the bottleneck for generation fluency. Understanding this bifurcation is key to seeing why deployment-side optimizations all revolve around "compressing decode overhead" — batch inference, continuous batching, KV Cache quantization are all filling the decode-phase holes. See Deployment & Serving for details.
3. Decoding Strategies: From Greedy to Sampling
Decoding strategies determine "how to pick tokens from the probability distribution." Start with two basic approaches:
1. Greedy Decoding
Directly pick the highest-probability token at each step:
text
next_token = argmax( probs )- Pros: deterministic, stable, reproducible; suitable for math/code/factual QA.
- Cons: short-sighted — single-step optimal ≠ globally optimal; prone to repetition (same word output repeatedly); no diversity, same question always same answer.
2. Temperature Sampling
First divide logits by temperature T, then softmax to probability, then sample randomly per probability:
text
p_i = exp( z_i / T ) / Σ_j exp( z_j / T )
T = 0 → degenerates to argmax (greedy)
T < 1 → probability distribution sharper, more conservative, more certain
T = 1 → preserves original distribution
T > 1 → distribution smoother, low-probability words "lifted up," more diverse, riskierTemperature isn't as simple as a "creativity knob"
Temperature adjusts "the smoothness of the distribution," not "smartness." Too high T lets the model pick wrong words, directly dropping scores on factual tasks; T=0 is most stable for math problems. "Creativity comes from "giving low-probability paths a bit of a chance" — but too many chances also bring hallucination and grammar errors (see Hallucination: Causes & Mitigation).
A practical intuition: temperature determines "willingness to take low-probability paths," top-p determines "how wide a candidate pool to allow." They act on different dimensions but both "relax/tighten" the distribution: T from 0 upward, distribution flattens; top-p from 1 downward, candidate pool narrows. In practice, first set temperature (style), then use top-p as a safety net (prevent drifting).
3. Top-k Truncation
Only sample from the top k highest-probability tokens (k commonly 10–100):
text
Candidate set = top k tokens by probability → normalize and samplePros: chops off long-tail low-probability words; cons: k is a global fixed value — in "sharp distribution" contexts (many words have near-max probability), too-small k actually discards reasonable options.
4. Top-p (Nucleus Sampling)
Dynamically select the smallest token set whose cumulative probability just exceeds p (p commonly 0.9–0.95):
text
1. Sort tokens by probability, high to low
2. Accumulate probability from high to low until cumulative ≥ p
3. Only normalize and sample within the "selected core set"Top-p is more adaptive than top-k: the set is smaller in sharp distributions, larger in flat ones. It's one of the default configurations for current APIs and open-source inference frameworks.
5. Min-p Truncation
Min-p (proposed 2024, a variant of nucleus sampling) filters by relative threshold: only keep tokens whose probability is not lower than "current max probability × min_p":
text
Keep condition: p_i ≥ min_p × max_prob
Example: max_prob = 0.6, min_p = 0.05 → keep tokens with prob ≥ 0.03Compared to top-p, min-p's filtering intensity changes dynamically with the distribution, typically balancing diversity and coherence better during long-text generation (in practice commonly used with temperature).
6. Beam Search: Maintaining Multiple Candidates
Beam search no longer "pick one at each step"; instead, it maintains beam_size candidate sequences at all times, extending all options for each candidate at each step, keeping the top beam_size by score, finally picking the sequence with the highest total score.
- Pros: can "see" longer-term combinations; commonly used in machine translation/summarization where deterministic optimality is pursued.
- Cons: limited quality gain on high-quality decoding decoders; also prone to repetition; compute cost scales linearly with beam_size; incompatible with streaming output (must wait for all candidates to finish).
Decoding Strategy Comparison Summary
| Strategy | Mechanism | Diversity | Determinism | Applicable scenario |
|---|---|---|---|---|
| Greedy | argmax per step | None | Fully deterministic | Math, code, factual QA |
| Temperature sampling | Divide by T then sample per probability | Higher with ↑T | T=0 deterministic | General generation, creative writing |
| Top-k | Truncate to top k then sample | Medium | Random | Quick long-tail pruning |
| Top-p | Sample within set where cumulative ≥ p | Medium-high | Random | General default (~0.9–0.95) |
| Min-p | Filter by relative max-prob threshold | Medium-high | Random | Long text, pursuit of coherence + diversity |
| Beam search | Keep beam candidates, pick best | Low | Approximately deterministic | Summarization, translation, constrained generation |
Modern practice combo
Mainstream API and inference framework recommendation: use temperature + top_p (or min-p) together; don't stack multiple truncation strategies (over-compresses the distribution). Common starting point: general dialogue temperature=0.7–0.9, top_p=0.9; code/math temperature=0–0.3; creative writing temperature≈1.0. Exact values vary by model; above are empirical starting points.
7. Stopping Conditions: When to Stop Generating
The autoregressive loop needs an explicit stopping condition; three common ones:
| Condition | Mechanism | Note |
|---|---|---|
| EOS token | Model outputs end-of-sequence token → stop | Natural finish, but may stop too early/late |
| max_tokens | Hard length cap | Prevents runaway, but may truncate sentences |
| Stop sequences | Stop when output matches specified string | Common for dialogue/code completion (e.g., delimiters, end markers) |
Production systems typically stack all three: max_tokens as fallback, stop sequences controlling format boundaries, EOS for natural finish. Poor stopping conditions directly affect output quality — e.g., in code completion, not setting a stop sequence may cause the model to generate irrelevant content after the function body.
4. Repetition Penalty: Controlling the Repeater
Temperature adjusts diversity but doesn't control "repetition" — the model may fall into loops of the same phrase or sentence pattern. Repetition penalty suppresses repetition by lowering the probability of "already appeared" tokens; two common types:
| Parameter | Mechanism | Typical range (OpenAI-style) |
|---|---|---|
| Frequency penalty | Linear penalty based on token appearance count | 0–2, commonly 0–1 |
| Presence penalty | Penalty once a token has appeared, regardless of count | 0–2, commonly 0–0.5 |
text
For already-appeared token t:
logit(t) -= frequency_penalty × count(t) # the more it appears, the heavier the penalty
logit(t) -= presence_penalty # appears once, penalized oncePenalties are a double-edged sword
Too-high penalties make the model "go around" — using obscure expressions to avoid repetition, resulting in awkward phrasing and degraded facts. First troubleshoot "should you change decoding strategy (e.g., add top-p) or whether the prompt itself lacks guidance" before reaching for penalties; penalties are the last resort. More thorough repeater governance returns to prompt engineering and the training side.
5. Dialogue Format: System Prompts and Message Templates
API inference typically takes structured messages as input: three roles — system, user, assistant. The system prompt is the highest-priority instruction area, setting role, behavioral boundaries, and output requirements:
text
system: You are a senior frontend engineer; answer technical questions with runnable example code.
user: How do you pass functions as props in Vue?
assistant: (model generates)- Put "long-term rules" in the system prompt: role, language, safety boundaries, output format; the earlier and clearer, the better the model follows.
- User/assistant alternate to form dialogue history; the model references history to maintain multi-turn consistency.
- Dialogue templates are each model family's private convention (ChatML, Llama templates, etc.); inference frameworks auto-apply based on the model's
chat_template. Why different models have different templates and what happens when templates are wrong relates to the boundary token design in Tokenization & Vocabulary.
Combining sampling params with dialogue systems
In dialogue products, the system prompt handles "stable behavioral constraints," sampling parameters handle "style control": treat temperature as a product feature (lower temperature for formal assistants, higher for creative ones), and let users adjust it on the frontend.
6. Streaming Output: Return as You Generate
Autoregressive generation is inherently "step-by-step producing," so it can stream (via SSE — Server-Sent Events or similar protocols): push each generated token to the client as it's generated, instead of waiting for full generation. Key metrics:
| Metric | Meaning | Relationship to streaming |
|---|---|---|
| TTFT (Time to First Token) | Time from request to first token return | Determined by prefill phase + first decode step |
| TPOT (Time per Output Token) | Interval between each generated token | Determined by decode-phase single-step latency |
| Throughput | Tokens generated per second | Target of batching and KV Cache optimization |
Product experience: streaming lets users "see the thinking process," dramatically reducing perceived latency (TTFT can reach hundreds of milliseconds); engineering-wise it pairs with continuous batching to boost GPU utilization — see Deployment & Serving.
Streaming and beam search are incompatible
Beam search explores multiple candidate sequences; it can't "output as it computes." Products needing streaming experience (dialogue, code completion) should use sampling-style strategies. This is one reason modern dialogue/coding products nearly all use sampling.
7. Sampling Hyperparameter Practical Advice Table
| Scenario | temperature | top_p | Penalties | Note |
|---|---|---|---|---|
| Math/code generation | 0–0.3 | 0.9–1.0 | Low or 0 | Want deterministic and correct; don't add diversity |
| Factual QA/customer service | 0.2–0.5 | 0.9 | 0.3–0.5 presence | Suppress repetitive boilerplate |
| General dialogue/summarization | 0.7–0.9 | 0.9 | 0–0.3 | Balance naturalness and stability |
| Creative writing/brainstorming | 0.9–1.2 | 0.95–1.0 | 0.5–1.0 frequency | Diversity first; accept hallucination risk |
| Code completion (IDE) | 0.1–0.3 | 0.9 | Low | Long completion paired with min-p recommended |
Mainstream API parameter mapping: OpenAI-style uses temperature (0–2), top_p, frequency_penalty, presence_penalty; Anthropic additionally exposes top_k; local inference frameworks (vLLM, llama.cpp) directly expose the full sampler combo. Note differences across vendors in parameter names and value ranges — read docs before migrating calls, don't assume 0–2 temperature range maps to 0–1.
How to tune parameters scientifically
Don't guess from intuition: fix model and prompt, change only one parameter, compare output quality on a fixed test set (optionally pair with LLM-as-a-judge or human blind A/B). Sampling parameters are part of product experience — worth versioning and regression-testing like UI. Prompt-side effects are in Prompt Engineering.
Further Reading
- Transformer Architecture Deep Dive — attention mechanism and QKV, the source of KV Cache
- Context & Long Context — long context's challenge to KV Cache and attention
- Prompt Engineering — how system prompts and output format affect generation
- Deployment & Serving — KV Cache, quantization, batching, and memory estimation
- Tokenization & Vocabulary — how token granularity affects generation speed and cost
- Prompting Practice — tuning sampling params together with prompts
References
- Holtzman et al. The Curious Case of Neural Text Degeneration (ICLR 2020, nucleus sampling) — the original top-p sampling paper, including repetition analysis
- Nguyen et al. Truncate or Penalize? A New Approach to Penalizing and Truncating the Next Token Distribution (2024) — the min-p truncation method
- Vaswani et al. Attention Is All You Need (NeurIPS 2017) — the original source of attention mechanism and QKV
- OpenAI API Documentation: Text generation (temperature/top_p/penalty parameter notes) — official definition and ranges for sampling parameters
- Kwon et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM, SOSP 2023) — KV Cache management and representative system for inference optimization