Skip to content

Inference Fundamentals: Autoregression and Sampling

At a glance The inference stage determines how the model "speaks." This article breaks down the autoregressive generation loop, the prefill/decode two phases of KV Cache, compares greedy decoding, temperature sampling, top-k, top-p, min-p, and beam search decoding strategies, and gives practical advice on temperature/top_p/penalties for different scenarios.

Inference Fundamentals: Autoregression and Sampling ​

Inference is the stage where the model truly "speaks" after training: given input, generating output token by token. Training determines "what the model knows"; inference configuration determines "how it says it" — the same model, temperature at 0 is a rigorous engineer, at 1.0 is a creative writer loose its constraints. That's the subject of this article.

This article unfolds along the path of "generation loop → KV Cache → decoding strategies → dialogue format → streaming output." Performance engineering for inference (quantization, batching, memory) is in Deployment & Serving; measuring sampling effects is in Evaluation & Benchmarks.

1. Autoregression: One Token at a Time ​

Most modern LLMs (GPT, Llama, Qwen, etc.) are autoregressive decoders: predict one next token at a time, append the new token to the input tail, then predict the next, looping until generation ends. This is the direct manifestation of language modeling's "predict the next word" paradigm at inference time.

text
Autoregressive generation loop (pseudocode):

Input tokens = [t1, t2, ..., tn]
Output tokens = []
while True:
    # 1. Forward pass: get the next-token probability distribution
    probs = model(tokens)          # shape [vocab_size]
    # 2. Pick a token per decoding strategy (greedy/sampling, see Section 3)
    next_token = decode(probs)
    # 3. Append to sequence
    tokens.append(next_token)
    outputs.append(next_token)
    # 4. Stop at EOS or max_tokens
    if next_token == EOS or len(outputs) >= max_tokens:
        break

Key implication: generating n tokens requires n forward passes; step t must "see" all preceding t-1 tokens. This brings two engineering consequences — KV Cache (save repeated computation) and streaming output (send as you compute, no need to wait for full generation).

Why autoregression is the de facto standard

Autoregression reduces "generating text" to "repeatedly doing what language models do best (predict the next word)," aligning the training objective perfectly with the generation objective (by contrast, BERT-style masked models need extra design for generation, see Transformer Architecture Deep Dive). The cost is generation can't be parallelized — this is the root problem KV Cache and batching optimizations address.

2. KV Cache: Prefill and Decode Two Phases ​

In the attention mechanism, every token needs to compute Query, Key, Value (QKV). At step t, the new token must compute attention with all preceding tokens' Key/Value. Without caching, step t would need to recompute K/V for the preceding t-1 tokens — redundant work.

KV Cache caches the Key and Value of all preceding tokens in GPU memory, computing only the new token's K/V during generation, directly reusing the cache:

text
Prefill phase (prompt phase):
  Process the full input prompt all at once, compute and cache K/V for all prompt tokens
  ── High parallelism, large throughput

Decode phase:
  Generate token by token, each step only needs:
    ① Compute the new token's Q/K/V
    ② Compute attention between the new token's K/V and cached K/V
    ③ Append new K/V to cache
  ── Serial, latency-dominated (delay determined by single-step time)

KV Cache is one of the main sources of memory overhead at inference time (in long context, it can even exceed weight memory), and is a key object in Context & Long Context and deployment optimization. Memory estimation, PagedAttention, etc. optimizations are in Deployment & Serving.

Long context ≠ unlimited KV Cache

KV Cache size grows linearly with sequence length; at long context (100K token range), memory pressure spikes dramatically. This is the gap between "model supports 128K context" and "can really run full 128K" — model capability, memory budget, and deployment config jointly determine the actual usable length.

The prefill and decode two phases have completely different meanings for user experience: prefill determines TTFT (first-token latency) — processing the full prompt all at once, short prompt fast, long prompt slow; decode determines TPOT (subsequent token interval) — token-by-token serial, the bottleneck for generation fluency. Understanding this bifurcation is key to seeing why deployment-side optimizations all revolve around "compressing decode overhead" — batch inference, continuous batching, KV Cache quantization are all filling the decode-phase holes. See Deployment & Serving for details.

3. Decoding Strategies: From Greedy to Sampling ​

Decoding strategies determine "how to pick tokens from the probability distribution." Start with two basic approaches:

1. Greedy Decoding ​

Directly pick the highest-probability token at each step:

text
next_token = argmax( probs )
  • Pros: deterministic, stable, reproducible; suitable for math/code/factual QA.
  • Cons: short-sighted — single-step optimal ≠ globally optimal; prone to repetition (same word output repeatedly); no diversity, same question always same answer.

2. Temperature Sampling ​

First divide logits by temperature T, then softmax to probability, then sample randomly per probability:

text
p_i = exp( z_i / T ) / Σ_j exp( z_j / T )

T = 0    → degenerates to argmax (greedy)
T < 1    → probability distribution sharper, more conservative, more certain
T = 1    → preserves original distribution
T > 1    → distribution smoother, low-probability words "lifted up," more diverse, riskier

Temperature isn't as simple as a "creativity knob"

Temperature adjusts "the smoothness of the distribution," not "smartness." Too high T lets the model pick wrong words, directly dropping scores on factual tasks; T=0 is most stable for math problems. "Creativity comes from "giving low-probability paths a bit of a chance" — but too many chances also bring hallucination and grammar errors (see Hallucination: Causes & Mitigation).

A practical intuition: temperature determines "willingness to take low-probability paths," top-p determines "how wide a candidate pool to allow." They act on different dimensions but both "relax/tighten" the distribution: T from 0 upward, distribution flattens; top-p from 1 downward, candidate pool narrows. In practice, first set temperature (style), then use top-p as a safety net (prevent drifting).

3. Top-k Truncation ​

Only sample from the top k highest-probability tokens (k commonly 10–100):

text
Candidate set = top k tokens by probability → normalize and sample

Pros: chops off long-tail low-probability words; cons: k is a global fixed value — in "sharp distribution" contexts (many words have near-max probability), too-small k actually discards reasonable options.

4. Top-p (Nucleus Sampling) ​

Dynamically select the smallest token set whose cumulative probability just exceeds p (p commonly 0.9–0.95):

text
1. Sort tokens by probability, high to low
2. Accumulate probability from high to low until cumulative ≥ p
3. Only normalize and sample within the "selected core set"

Top-p is more adaptive than top-k: the set is smaller in sharp distributions, larger in flat ones. It's one of the default configurations for current APIs and open-source inference frameworks.

5. Min-p Truncation ​

Min-p (proposed 2024, a variant of nucleus sampling) filters by relative threshold: only keep tokens whose probability is not lower than "current max probability × min_p":

text
Keep condition: p_i ≥ min_p × max_prob
Example: max_prob = 0.6, min_p = 0.05 → keep tokens with prob ≥ 0.03

Compared to top-p, min-p's filtering intensity changes dynamically with the distribution, typically balancing diversity and coherence better during long-text generation (in practice commonly used with temperature).

6. Beam Search: Maintaining Multiple Candidates ​

Beam search no longer "pick one at each step"; instead, it maintains beam_size candidate sequences at all times, extending all options for each candidate at each step, keeping the top beam_size by score, finally picking the sequence with the highest total score.

  • Pros: can "see" longer-term combinations; commonly used in machine translation/summarization where deterministic optimality is pursued.
  • Cons: limited quality gain on high-quality decoding decoders; also prone to repetition; compute cost scales linearly with beam_size; incompatible with streaming output (must wait for all candidates to finish).

Decoding Strategy Comparison Summary ​

StrategyMechanismDiversityDeterminismApplicable scenario
Greedyargmax per stepNoneFully deterministicMath, code, factual QA
Temperature samplingDivide by T then sample per probabilityHigher with ↑TT=0 deterministicGeneral generation, creative writing
Top-kTruncate to top k then sampleMediumRandomQuick long-tail pruning
Top-pSample within set where cumulative ≥ pMedium-highRandomGeneral default (~0.9–0.95)
Min-pFilter by relative max-prob thresholdMedium-highRandomLong text, pursuit of coherence + diversity
Beam searchKeep beam candidates, pick bestLowApproximately deterministicSummarization, translation, constrained generation

Modern practice combo

Mainstream API and inference framework recommendation: use temperature + top_p (or min-p) together; don't stack multiple truncation strategies (over-compresses the distribution). Common starting point: general dialogue temperature=0.7–0.9, top_p=0.9; code/math temperature=0–0.3; creative writing temperature≈1.0. Exact values vary by model; above are empirical starting points.

7. Stopping Conditions: When to Stop Generating ​

The autoregressive loop needs an explicit stopping condition; three common ones:

ConditionMechanismNote
EOS tokenModel outputs end-of-sequence token → stopNatural finish, but may stop too early/late
max_tokensHard length capPrevents runaway, but may truncate sentences
Stop sequencesStop when output matches specified stringCommon for dialogue/code completion (e.g., delimiters, end markers)

Production systems typically stack all three: max_tokens as fallback, stop sequences controlling format boundaries, EOS for natural finish. Poor stopping conditions directly affect output quality — e.g., in code completion, not setting a stop sequence may cause the model to generate irrelevant content after the function body.

4. Repetition Penalty: Controlling the Repeater ​

Temperature adjusts diversity but doesn't control "repetition" — the model may fall into loops of the same phrase or sentence pattern. Repetition penalty suppresses repetition by lowering the probability of "already appeared" tokens; two common types:

ParameterMechanismTypical range (OpenAI-style)
Frequency penaltyLinear penalty based on token appearance count0–2, commonly 0–1
Presence penaltyPenalty once a token has appeared, regardless of count0–2, commonly 0–0.5
text
For already-appeared token t:
  logit(t) -= frequency_penalty × count(t)     # the more it appears, the heavier the penalty
  logit(t) -= presence_penalty                 # appears once, penalized once

Penalties are a double-edged sword

Too-high penalties make the model "go around" — using obscure expressions to avoid repetition, resulting in awkward phrasing and degraded facts. First troubleshoot "should you change decoding strategy (e.g., add top-p) or whether the prompt itself lacks guidance" before reaching for penalties; penalties are the last resort. More thorough repeater governance returns to prompt engineering and the training side.

5. Dialogue Format: System Prompts and Message Templates ​

API inference typically takes structured messages as input: three roles — system, user, assistant. The system prompt is the highest-priority instruction area, setting role, behavioral boundaries, and output requirements:

text
system:  You are a senior frontend engineer; answer technical questions with runnable example code.
user:    How do you pass functions as props in Vue?
assistant: (model generates)
  • Put "long-term rules" in the system prompt: role, language, safety boundaries, output format; the earlier and clearer, the better the model follows.
  • User/assistant alternate to form dialogue history; the model references history to maintain multi-turn consistency.
  • Dialogue templates are each model family's private convention (ChatML, Llama templates, etc.); inference frameworks auto-apply based on the model's chat_template. Why different models have different templates and what happens when templates are wrong relates to the boundary token design in Tokenization & Vocabulary.

Combining sampling params with dialogue systems

In dialogue products, the system prompt handles "stable behavioral constraints," sampling parameters handle "style control": treat temperature as a product feature (lower temperature for formal assistants, higher for creative ones), and let users adjust it on the frontend.

6. Streaming Output: Return as You Generate ​

Autoregressive generation is inherently "step-by-step producing," so it can stream (via SSE — Server-Sent Events or similar protocols): push each generated token to the client as it's generated, instead of waiting for full generation. Key metrics:

MetricMeaningRelationship to streaming
TTFT (Time to First Token)Time from request to first token returnDetermined by prefill phase + first decode step
TPOT (Time per Output Token)Interval between each generated tokenDetermined by decode-phase single-step latency
ThroughputTokens generated per secondTarget of batching and KV Cache optimization

Product experience: streaming lets users "see the thinking process," dramatically reducing perceived latency (TTFT can reach hundreds of milliseconds); engineering-wise it pairs with continuous batching to boost GPU utilization — see Deployment & Serving.

Streaming and beam search are incompatible

Beam search explores multiple candidate sequences; it can't "output as it computes." Products needing streaming experience (dialogue, code completion) should use sampling-style strategies. This is one reason modern dialogue/coding products nearly all use sampling.

7. Sampling Hyperparameter Practical Advice Table ​

Scenariotemperaturetop_pPenaltiesNote
Math/code generation0–0.30.9–1.0Low or 0Want deterministic and correct; don't add diversity
Factual QA/customer service0.2–0.50.90.3–0.5 presenceSuppress repetitive boilerplate
General dialogue/summarization0.7–0.90.90–0.3Balance naturalness and stability
Creative writing/brainstorming0.9–1.20.95–1.00.5–1.0 frequencyDiversity first; accept hallucination risk
Code completion (IDE)0.1–0.30.9LowLong completion paired with min-p recommended

Mainstream API parameter mapping: OpenAI-style uses temperature (0–2), top_p, frequency_penalty, presence_penalty; Anthropic additionally exposes top_k; local inference frameworks (vLLM, llama.cpp) directly expose the full sampler combo. Note differences across vendors in parameter names and value ranges — read docs before migrating calls, don't assume 0–2 temperature range maps to 0–1.

How to tune parameters scientifically

Don't guess from intuition: fix model and prompt, change only one parameter, compare output quality on a fixed test set (optionally pair with LLM-as-a-judge or human blind A/B). Sampling parameters are part of product experience — worth versioning and regression-testing like UI. Prompt-side effects are in Prompt Engineering.

Further Reading ​

References ​