Skip to content

Language Modeling: The Next-Token Prediction Paradigm

At a glance Language modeling is the foundation of all large language models: treating 'predict the next token' as a proxy objective to estimate the probability distribution of text sequences. This article clarifies the inheritance chain from N-gram to neural language models to pretraining paradigms, conditional probability decomposition, perplexity and cross-entropy, and 'why predicting the next token learns knowledge.'

Language Modeling: The Next-Token Prediction Paradigm ​

The goal of language modeling is to learn to assign probabilities to text: given preceding context, estimate the probability distribution of the next token (or word). It is a classic task dating back to the statistical NLP era, and in the 2020s it became the sole core training objective for large language models (LLMs). This seemingly simple proxy task of "predicting the next token" has become the source of knowledge, reasoning, and emergent abilities — powered by data and scale. Understanding language modeling is like holding the master key to understanding all LLM phenomena.

One-line summary: an LLM is a probability model that excels at "next-token prediction." For the full background and capability overview, see What Are Large Language Models?.

1. Formalization: What Probability Does a Language Model Compute ​

A language model estimates the joint probability $$P(w_1, w_2, \dots, w_n)$$ of a word sequence $$w_1, w_2, \dots, w_n$$. Estimating the joint distribution directly in high-dimensional space is infeasible, so we use the chain rule (conditional probability decomposition) to factor it into a product of conditional probabilities:

text
P(w1, w2, ..., wn) = P(w1) · P(w2|w1) · P(w3|w1,w2) · ... · P(wn|w1,...,w(n-1))

                    = ∏_{t=1}^{n} P(wt | w1, w2, ..., w(t-1))

In other words: "scoring a piece of text is equivalent to predicting the next token at each position, then multiplying all the prediction probabilities together." Conversely, starting from the conditional probability $$P(w_t \mid w_{<t})$$ and sampling step by step gives us generation — which is why "language models can naturally write text": writing an essay is essentially a repeated process of next-token prediction, feeding predicted results back into the input. This is the foundation for all decoding strategies discussed in Inference Fundamentals: Autoregression and Sampling.

This decomposition has three direct implications: scoring and generation are the same thing (scoring multiplies probabilities position by position, generation samples token by token); training data needs no additional labels (any text natively carries "preceding context → next token" supervision pairs); and the longer the sequence, the smaller the joint probability (more multiplicative terms cause exponential decay). In practice, we compare average loss (NLL) rather than raw joint probabilities.

Why "next token" and not "next sentence"?

The finer the prediction granularity, the denser the signal, and the higher the training efficiency. At the word/subword level, each training text can contribute hundreds or thousands of independent supervision signals, forcing the model to "bet" on the entire vocabulary at every position — this dense supervision is the key that makes large-scale training feasible. (Some researchers have explored alternatives like "predicting the next sentence" or "contrastive learning," but no objective has truly displaced next-token prediction for general capabilities — it nearly free provides dense signals, is natively parallelizable, and aligns perfectly with generation tasks.)

1. N-gram Language Models and the Markov Assumption ​

The classic approach from the statistical era is the N-gram model: it introduces the Markov assumption — the next word depends only on the most recent $$n-1$$ words, independent of earlier history:

text
P(wt | w1, ..., w(t-1)) ≈ P(wt | w(t-n+1), ..., w(t-1))

Parameters are estimated directly from frequency counts in the corpus (maximum likelihood): $$P(w_t \mid w_{t-n+1:t-1}) = \text{count}(w_{t-n+1:t}) / \text{count}(w_{t-n+1:t-1})$$. Larger n yields more accurate context but sparser counts, leading to a whole toolkit of remedial techniques: smoothing, backoff, and others. The essence of smoothing is "assigning a non-zero probability to unseen n-grams." Two mainstream approaches: backoff (fall back to shorter n-grams to borrow counts) and interpolation (weighted average of estimates across different orders). But none can eliminate sparsity — the number of n-grams grows exponentially with n, while corpus size is finite. This is the ceiling of the "estimate distributions by counting" approach.

The N-gram lesson profoundly influenced everything that followed: the fundamental contradiction that "longer context → sparser data" can only be mitigated by "representing longer history with fewer parameters" — and that's exactly what neural networks do.

2. Bengio 2003: Neural Probabilistic Language Model ​

Bengio et al.'s 2003 paper A Neural Probabilistic Language Model (JMLR) is the starting point for neural language models. It maps words to low-dimensional dense vectors (embeddings), then feeds them through a feedforward network with a hidden layer to produce a probability distribution over the next word:

text
Input: one-hot vectors for the previous n-1 words
  → Look up word embedding table C (dimension d, much smaller than vocab V)
  → Concatenate/sum to get context representation
  → One hidden layer with tanh activation
  → Softmax output P(wt | context) (distribution over vocab V)
Training objective: maximize log-likelihood of training corpus
  (equivalent to minimizing cross-entropy)

Two key contributions: first, using distributed representations to fight the curse of dimensionality, where meaning is carried by low-dimensional vectors rather than independent indices; second, the side effect — learned word vectors became the precursor to representation learning approaches like Word2Vec. There were also a few details that seemed minor at the time but had lasting impact: the softmax output layer costs O(V) (V is vocabulary size, potentially tens of thousands), which directly inspired acceleration techniques like hierarchical softmax, negative sampling, and sampled softmax; fixed window n locks usable context within the architecture; and "word representation + composition rule" are jointly learned under one objective — this "end-to-end" philosophy is the proto-form of the later deep learning paradigm. Its limitations (fixed window, single-layer nonlinearity) were soon surpassed by recurrent neural networks.

3. From RNN/LSTM to Transformer ​

RNN language models use a recurrent hidden state $$h_t = f(h_{t-1}, x_t)$$ to compress arbitrary-length prefixes into a single vector, theoretically breaking free of fixed windows; LSTM gates mitigate vanishing gradients and became the de facto standard for language modeling at one point. But RNNs have two fatal flaws: sequential, step-by-step computation prevents parallelism, and long-range dependencies are still diluted by "compression loss" — no matter how elegantly LSTM gates work, stuffing entire history into a fixed-dimensional vector is inherently lossy compression, with information from more distant positions decaying more severely. This is the watershed between the "compress to fixed length" approach and the "direct interconnections at each position" approach.

The 2017 Attention Is All You Need Transformer architecture deep dive replaced recurrence with self-attention that lets any two positions directly "see" each other, solving both parallelism and long-range dependency in one stroke — and also changing the shape of language modeling itself. A comparison of the three paradigms:

ParadigmContext ModelingParallelismLong-Range AbilityData ScaleRepresentative
N-gramFixed n-gram windowNatively parallelPoor (limited by n)10⁶~10⁸ wordsKenLM, etc.
RNN/LSTMCompressed into recurrent stateNone (step-by-step)Moderate (gradient/compression limited)10⁸~10⁹ wordsBest pre-GPT-2
TransformerFull-connect via self-attentionFully parallelStrong (O(n²) cost)10¹²~10¹³ tokensGPT, Llama, Qwen

2. Two Pretraining Objectives: Autoregressive vs Masked ​

"Predict the next token" is strictly an autoregressive (AR) objective: predict left-to-right, seeing only left context. This is natively aligned with generation. Another classic objective is the BERT-style masked language modeling (MLM): randomly mask about 15% of positions, and the model predicts the masked tokens using both left and right context — a "fill in the blank" approach rather than "continue writing." A comparison:

DimensionAutoregressive (GPT-style)Masked (BERT-style)
Prediction directionOne-directional (left only)Bidirectional (both sides)
Objective form$$P(w_t \mid w_{<t})$$$$P(w_{\text{mask}} \mid w_{\backslash \text{mask}})$$
Loss computationEvery position countsOnly masked positions count
StrengthsGeneration, continuation, few-shot reasoningUnderstanding, representations, retrieval/classification
Representative modelsGPT-1/2/3, Llama, QwenBERT, RoBERTa

Modern LLMs overwhelmingly follow the autoregressive route. Three reasons: first, generation is the harder objective — if you can generate, you can naturally also judge (via autoregressive scoring), but a "fill-in-the-blank" objective can't directly write long text; second, autoregressive objectives have every token participate in training, yielding high sample efficiency; third, few-shot/in-context learning emerges naturally under autoregressive objectives — GPT-3's few-shot phenomenon is the definitive evidence for this route.

To be clear, this does not negate the value of bidirectional encoding — the BERT route still leads in representation quality and retrieval/classification, and encoder-decoder models like T5 remain competitive on tasks that combine "input understanding + output generation" like translation and summarization. It's just that when the goal is "a universal open-domain assistant," the decoder-only autoregressive form unifies "modeling, generation, and few-shot learning" under one objective, keeping both engineering and theory cleaner. For the full story on both routes, see the GPT series and BERT and Encoder Families.

Why "pretraining" is a paradigm shift

Early neural language models were trained "from scratch on task-specific data"; GPT/BERT in 2018 established a two-phase paradigm: first pretrain on massive general-purpose corpora via language modeling, then (optionally) fine-tune. Language modeling went from being "one task" to "the mechanism for acquiring general capabilities." See Pretraining: Data and Objectives and Evolution Timeline.

3. Language Modeling as a "Unified Interface" ​

Another worth-memorizing perspective: whether it's GPT's autoregressive, BERT's masked, or T5's text-to-text, they can all be viewed as "give the model some text input, and it predicts some text output." Encoding translation, QA, summarization, and classification all as text-to-text is the fundamental reason LLMs can "handle every task with one model" — this is the deepest legacy of the "predict the next token" paradigm: language itself becomes the only interface between tasks.

3. Training Objectives: Cross-Entropy and Perplexity ​

1. Cross-Entropy Loss ​

Training maximizes the likelihood of the training text, which is equivalent to minimizing cross-entropy. For each position, the model outputs a probability distribution $$q_t$$ over the vocabulary, and the true word corresponds to a one-hot distribution $$p_t$$:

text
Per-position loss: L_t = -Σ_w p_t(w) · log q_t(w) = -log q_t(w_t)
Sequence loss (average negative log-likelihood, NLL):
L = -(1/n) Σ_t log q_t(w_t)

Intuitively: the closer the model's probability for the true word is to 1, the smaller the loss; assigning high probability to a wrong word gets harshly penalized. Cross-entropy gradients are proportional to "predicted probability − true distribution" — larger errors drive bigger pushes, leading to stable convergence. That's why it's the default loss for all language models.

2. Perplexity (PPL): The Built-in Metric Every Language Model Has ​

Perplexity (PPL) = exp(average negative log-likelihood):

text
PPL = exp(L) = exp( -(1/n) Σ_t log q_t(w_t) )

Intuitive reading: perplexity ≈ the average number of candidate tokens the model has to "guess between" at each step. A uniform six-sided die has perplexity 6; a fair binary coin has 2. Lower PPL means the model's predictions are more certain — it "gets" the text better. A few properties to note:

Calculate it yourself

Suppose the vocabulary has only two words, A and B, the test sentence is "A A B A", and the model's predicted probabilities are 0.8, 0.9, 0.7, 0.6. Average NLL = -(ln0.8 + ln0.9 + ln0.7 + ln0.6)/4 ≈ 0.34, PPL = e^0.34 ≈ 1.40 — the model hesitates between fewer than 1.5 candidates on average, quite confident. If at one position the probability is only 0.2, PPL shoots up immediately: low-probability positions contribute non-linearly to perplexity — this is why PPL readily exposes models that "are generally fluent but occasionally make severe errors."

PropertyNote
Dimensionless, comparableCan directly compare models side-by-side under the same tokenization
Strongly tied to tokenizationLarger vocabularies, finer token splits → typically lower PPL numbers (see Tokenization & Vocab); cross-tokenization comparisons are invalid
Training loss = validation PPL"Watching the loss curve" in pretraining is really watching validation PPL
Doesn't perfectly align with human qualityA 0.1 PPL drop might mean noticeably better generation quality, or might just mean the model is "more conservative"

3. How to Use PPL: An Engineering Perspective ​

PPL is the most common "thermometer" in pretraining, but discipline is needed:

  • Only compare under the same tokenization: before comparing PPL across models, confirm the tokenizer matches (see Tokenization & Vocab); otherwise, number differences come from tokenization rather than model capability.
  • Use a matched-distribution eval set: the gap between training distribution and target distribution shows up directly in PPL; the eval set should match the deployment scenario.
  • Focus on relative changes, not absolute values: the slope of PPL decline during training, and PPL differences across data mixes, carry more information than any single number.
  • Non-linearity between PPL and quality: a 0.1 PPL drop may mean noticeably better quality, or just more conservative behavior. The most economical approach: run a small batch of human evaluation to calibrate.

Low perplexity ≠ no hallucination

PPL measures "how well the model fits the training distribution," not "whether it's right." A model can assign very low PPL to fluent fabricated content — hallucination often happens precisely when the model is "confidently making things up." PPL is the starting point of Evaluation & Benchmarks, far from the endpoint. See Hallucination: Causes and Mitigation for the full picture.

4. Why "Predicting the Next Token" Learns Knowledge ​

A natural question: how can merely "guessing words" lead to learning knowledge, reasoning, and even writing code? Three perspectives help:

  1. Compression is understanding. Predicting the next token is essentially "losslessly compressing text." To predict well, the model must learn the text's generative mechanism — grammar, common sense, factual associations, causal structures. A model that can compress encyclopedic content to a small size has necessarily internalized the knowledge within. This aligns with the information-theoretic view that "understanding = effective compression."

  2. Co-occurrence statistics carry knowledge. Factual knowledge largely manifests as predictable co-occurrence relationships: the "Paris is the capital of ___" blank-filling task requires encoding the "Paris → France" connection across vast amounts of text. The model doesn't retrieve stored knowledge — it "kneads" billions of co-occurrence statistics into parameters.

  3. Scale turns "pattern memorization" into "general capability". Small models only learn surface collocations; when parameters and data cross certain scale thresholds, the model starts generalizing scattered patterns into composable rules (syntax generation, multi-step reasoning, instruction following). This qualitative leap from quantitative accumulation is the subject of Scaling Laws and "emergent abilities."

Empirical evidence accumulates yearly: GPT-3 showed that few-shot prompting — a skill never explicitly trained — emerges automatically once scale crosses a certain magnitude; InstructGPT/ChatGPT demonstrated that "understanding instructions" can be bolted on through post-training, while the underlying language and knowledge base still come from pretraining next-token prediction; more compelling evidence comes from knowledge localization — researchers using probing techniques can locate specific facts (e.g., "the Eiffel Tower is in Paris") within attention heads and FFN layers, proving that knowledge indeed accumulates as parameters under this objective. In short, next-token prediction is not an excuse for "memorization" — it's a mechanism for compressing a world model. Whether the model truly "understands" depends on whether scale is large enough and data diverse enough.

An analogy often drawn is language acquisition: children exposed to vast amounts of language develop the ability to "construct their own sentences" from "mimicking adults" — they learn not individual sentences but the generative rules underneath. Language models digitize the same phenomenon via "predicting the next token": rules (parameters) emerge from billions of "prediction exercises." Of course, the analogy has limits: the model has no body, no intention; its learned "rules" are the projection of the world in text, not lived experience — which is one of the deep roots of hallucination and safety problems.

One-line summary of this page

Language modeling = conditional probability decomposition of text sequences + per-position cross-entropy training; PPL is its built-in measure; "predicting the next token" is the origin of all LLM capabilities — generation, knowledge, in-context learning all grow from this single objective.

5. Trade-offs and Boundaries ​

  • The ceiling of probabilistic models: language modeling learns "what text looks like," not "what the world is." It excels at in-distribution fluency and plausibility, but does not guarantee factual or logical correctness.
  • Objective–task misalignment: the pretraining objective (generalist) and application objectives (QA, translation, Agent) diverge — fine-tuning and alignment bridge the gap.
  • The evaluation dilemma: PPL is automatically computable but semantically limited; downstream capabilities rely on benchmarks, but benchmarks themselves have saturation and contamination issues (see Evaluation & Benchmarks for the full framework).
  • Generation cost: each autoregressive generation step requires a fresh forward pass — understanding this cost structure is central to Inference Fundamentals.

Language Models vs. Human Language Ability ​

Placing "next-token prediction" alongside human abilities clarifies the boundaries of this approach:

AbilityLanguage Model (Next-Token Paradigm)Humans
Vocabulary & SyntaxExtremely strong in-distribution, fluent generationHigh learning cost, but generalizes well
World KnowledgeDepends on corpus coverage & co-occurrence statsGained through experience & reasoning
Long-range PlanningWeak: greedy per-token, lacks global objectivesStrong: has goals and plans
Fact-CheckingNo proactive verification, "fluency-first" tendencyQuestions and verifies
Knowledge UpdatesRequires retraining or external retrieval (see RAG)Lifelong learning

6. From Language Modeling to Products: How Capabilities Emerge ​

Linking the knowledge from this page into a product chain:

text
Next-Token Prediction (this page)
  → Few-shot/In-Context Learning (GPT-3: given a few examples, it can do the task)
  → Instruction Following (SFT/Alignment: turning "can predict" into "can follow", see fine-tuning & alignment)
  → Reasoning Enhancement (CoT, Chain-of-Thought, see prompting)
  → Tools & Memory (RAG, Agent: giving the model external retrieval and execution)
  → Multimodal (adding images/audio beyond text)

Each rung doesn't replace language modeling — it adds a capability layer on top of the language modeling base. Understanding this hierarchy clarifies why "almost every breakthrough in foundation models starts with better language modeling."

Open Questions in Language Modeling ​

Even as successful as "next-token prediction" is, it has three acknowledged shortcomings that remain active research priorities: long-range coherence and planning (per-token prediction natively lacks a "think the whole piece first, then write" global objective), factuality and consistency (in-distribution fluency ≠ factual correctness, see Hallucination: Causes and Mitigation), and objective-alignment misalignment (language modeling learns "what something looks like," products need "what it should do" — the bridge is built by alignment). A noteworthy trend: combining next-token prediction with other objectives (planning at inference, retrieval, multimodal alignment) is becoming the common direction for next-generation models.

Further Reading ​

References ​