Skip to content

Large Language Models (LLMs)

At a glance The core engine of the generative AI wave — how LLMs learn through next-token prediction, the three-stage training pipeline, emergent abilities, architecture families, leading models, hard limits, and the trade-offs between prompt, RAG, and agent approaches.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Large Language Models (LLMs) ​

Definition in One Sentence ​

A large language model (LLM) is a language model built on the Transformer architecture, pretrained on massive amounts of text with "predict the next token" as its learning objective, and usually packing billions of parameters or more.

It is both a "language" model and a "world" model: to predict the next word, the model is forced to distill grammar rules, factual knowledge, logical relations, even persona and tone, out of billions of web pages, books, and code. The Transformer paper landed in 2017, GPT and BERT followed in 2018, and ChatGPT detonated the mass market in late 2022. LLMs are now the core engine of the generative AI wave — most of the hot topics in What Are AI's Hot Concepts revolve around them.

Who uses LLMs? Practically everyone, just in different ways: ordinary people ask questions of chat products; developers build applications through APIs; researchers keep making them bigger, smaller, and better understood; enterprises use them to reshape customer service, writing, code, search, and nearly every other business function. It is no exaggeration to say that the LLM is the single most significant AI breakthrough of the past decade. To see where it sits on the wider technical map, start with The Concept Landscape and A Brief History of AI.

The Core Mechanism: Autoregressive Language Modeling ​

Everything an LLM can do boils down to one operation: given a piece of text, predict what the most likely next token is.

From Tokens to Conditional Probability ​

Raw text is first split into tokens — words, subwords, or characters. The model then frames "predict the next token" as a conditional probability:

text
P(xₜ | x₁, x₂, …, xₜ₋₁)

Given the first t-1 tokens, predict a probability distribution over the t-th token. The joint probability of the whole text is the product of every token's conditional probability:

text
P(x₁, x₂, …, xₙ) = ∏ₜ P(xₜ | x₁, …, xₜ₋₁)

Training maximizes this log-likelihood (equivalently, minimizes the cross-entropy loss):

text
L = -Σₜ log P(xₜ | x₁, …, xₜ₋₁)

Training and Inference: Two Sides of the Same Coin ​

The elegance of this mechanism is that at training time the model learns nothing but "continue the text," and at inference time that is exactly all it does — both share the same code path. During training, real text serves as the answer key for step-by-step prediction (teacher forcing); during inference, each token the model just generated becomes part of the input for the next step, and the whole output is produced autoregressively, one token at a time:

python
# Autoregressive generation: one token at a time
def generate(model, prompt, max_new_tokens=100):
    tokens = tokenize(prompt)
    for _ in range(max_new_tokens):
        logits = model(tokens)             # the model outputs a distribution over the next token
        next_token = sample(logits[-1])    # sample from the distribution (temperature/top-p controllable)
        tokens.append(next_token)
        if next_token == EOS:              # stop at the end-of-sequence token
            break
    return detokenize(tokens)

Key takeaway

An LLM is not a system that "looks up answers by keyword." It is a generator of probability distributions over a vocabulary. Every generation is a sample — the same question can get different answers on different runs. That is not a bug; it is an inherent property of a probabilistic model. Keep this in mind as the starting point for understanding every capability and flaw that follows, hallucination above all.

"Predict the next token" sounds simple, but for predictions to be accurate, the model must internalize vast amounts of knowledge. The Transformer's self-attention lets every token "see" every other token in the context, which is what makes modeling long-range dependencies possible — the physical foundation of LLMs. Details in Transformers and Attention.

The Three-Stage Training Paradigm: Pretraining → Post-Training ​

An LLM's abilities are not acquired in a single pass; they are cultivated in stages. Modern LLMs follow a "pretraining + post-training" paradigm, and post-training itself splits into supervised fine-tuning and alignment:

text
    Massive text corpora (trillions of tokens: web pages, books, code)
                          │
                          ▼
        ┌────────────────────────────────┐
        │  (1) Pretraining               │
        │  Goal: predict the next token  │
        │  Output: a base model that can │
        │  "continue text"               │
        └────────────────────────────────┘
                          │
                          ▼
        ┌────────────────────────────────┐
        │  (2) Supervised Fine-Tuning    │
        │      (SFT / instruction tuning)│
        │  Data: human-written           │
        │  instruction-answer pairs      │
        │  Output: organizes knowledge   │
        │  into proper "answers"         │
        └────────────────────────────────┘
                          │
                          ▼
        ┌────────────────────────────────┐
        │  (3) Alignment (RLHF / DPO)    │
        │  Data: human preference        │
        │  rankings over responses       │
        │  Output: a model that is safe, │
        │  honest, and helpful           │
        └────────────────────────────────┘
                          │
                          ▼
                 Deploy for inference
  • Pretraining: next-token prediction over massive unlabeled text. It takes weeks and tens of millions of dollars and produces a "base model." The base model already has language ability and a huge store of knowledge, but it "only continues text — it doesn't answer questions properly."
  • Supervised Fine-Tuning (SFT): continued training on tens of thousands of "instruction → desired answer" pairs, teaching the model to answer whatever the user asks. This is the step that turns a text-continuing infant into an assistant that answers questions; the mechanics are covered in Fine-Tuning and PEFT (LoRA).
  • Alignment: trained on human preference data (rankings over multiple responses) via RLHF or DPO, making the model readier to refuse harmful requests, more honest about what it doesn't know, and less sycophantic. Methods and costs in Alignment: RLHF and DPO.

One sentence to remember

Pretraining sets the ceiling on capability; post-training sets the behavioral style. Post-training cannot conjure knowledge the base model never learned — it can only shape existing abilities into what users want. Many "the model got dumber" complaints trace back to alignment suppressing the model's expressiveness too aggressively.

If you plan to reproduce this pipeline yourself (even at small scale), follow the hands-on route in Fine-Tune Your Own LLM.

Emergent Abilities: When Scale Changes Quality ​

Once model scale crosses a certain threshold (usually somewhere past tens of billions of parameters), abilities that were never explicitly required by the training objective suddenly appear. These are emergent abilities:

AbilityOne-line descriptionHow you trigger it
In-context learningLearns a new task from a few examples in the prompt, with no parameter updatesProvide 1–5 input-output examples in the prompt
Instruction followingUnderstands and executes natural-language instructionsDescribe the task in natural language
Chain-of-thought (CoT) reasoning"Think step by step, then answer" markedly improves math/logic accuracyAsk the model to "think step by step"
Few-shotHandles never-before-seen tasks from a handful of examplesEmbed examples in the prompt

These abilities are nothing like the classical "train one model per task" playbook, and they are precisely the watershed separating LLMs from every NLP model that came before. The research evidence comes from Wei et al. (2022), Emergent Abilities of Large Language Models — see Core Papers, Close Up.

How Emergence Relates to Scaling Laws ​

Scaling laws describe a predictable power-law relationship between "bigger model + more data + more compute" and falling loss: Kaplan et al. (2020) found that loss falls as a power law in parameters, data, and compute; Chinchilla (Hoffmann et al., 2022) then derived the "compute-optimal" recipe — when model parameters double, training tokens should double as well, which is where today's rule of thumb of "roughly 20 tokens per parameter" for training large models comes from.

The relationship between the two fits in one line:

The one-line verdict

Scaling laws deliver predictable, smooth improvement; emergence delivers unpredictable, discontinuous jumps. Scale pushes loss steadily downward, while certain abilities suddenly "unlock" at particular points on the loss curve — we cannot precisely predict at what scale they will appear, only stumble into them experimentally. That is the deep motive behind frontier teams still racing to pile on scale; see Frontier Progress.

This also explains why the 2020s "arms race" is so fierce: whoever reaches the next threshold scale first unlocks the next batch of abilities first.

Architecture Families: From BERT and GPT to T5 ​

Depending on how the Transformer is used, language models fall into three families:

FamilyRepresentative modelsAttentionPretraining objectiveGood atWhere they stand today
Encoder-onlyBERT, RoBERTaBidirectionalMasked language modeling (MLM)Understanding, representations, classification, retrievalThe workhorse for semantic vectors / embeddings
Decoder-onlyGPT family, Llama, Qwen, DeepSeekUnidirectional (causal)Next-token predictionGeneration, dialogue, reasoningThe absolute mainstream
Encoder-DecoderT5, BARTBidirectional encoding + autoregressive decodingSpan infillingTranslation, summarization, other transduction tasksStill valuable for specific tasks

Why Decoder-only Won in the End ​

In 2018, BERT (a bidirectional encoder) crushed GPT-1 on understanding tasks, and the industry briefly believed "encoders are the way." From 2023 on, virtually every flagship model has been Decoder-only. Three reasons:

  1. Training and inference objectives match: generation is the only way an LLM produces output, and the causal (unidirectional) next-token training objective is exactly isomorphic to autoregressive inference — no "encoder understands, decoder generates" objective mismatch.
  2. Engineering efficiency: causal attention is a natural fit for inference acceleration (KV cache reuse), and the Decoder-only architecture is simpler and easier to scale to enormous sizes.
  3. Task unification: translation, summarization, Q&A, coding — all of it can be expressed as text-in, text-out sequence-to-sequence tasks, which a single Decoder-only model handles without any per-task head designs.

Fun fact

T5's own conclusion ran opposite to today's consensus: it found encoder-decoder slightly better than decoder-only on transduction tasks. History showed that decoder-only models overtook it on the strength of scale and ecosystem (OpenAI's GPT line simply kept getting bigger). The final winner of an architecture debate is usually decided by engineering scalability and ecosystem, not single-benchmark accuracy.

For the internals — attention, positional encodings, KV cache, and more — see Transformers and Attention.

Key Models Compared ​

As of mid-2025, the global LLM landscape is roughly "three big US closed-source players plus a four-way open-source contest." The table below exists to build intuition — models update monthly; for the latest leaderboards, check Models and Leaderboards Quick Reference.

ModelMakerReleasedParametersOpen sourceNotes
GPT-4oOpenAI2024-05Undisclosed (believed MoE)NoNatively multimodal, low latency; ChatGPT's mainstay — see ChatGPT and Conversational AI
Claude 3.5 / 3.7 SonnetAnthropic2024-06 / 2025-02UndisclosedNoLong context, strong code and long-form writing, standout safety alignment
Gemini 1.5 Pro / 2.5 ProGoogle2024-02 / 2025-03UndisclosedNoMillion-token-scale context window, deep multimodal integration
Llama 3.1Meta2024-078B / 70B / 405BYesThe open-weights flagship; largest ecosystem
Qwen2.5Alibaba2024-090.5B–72BYesMultilingual with standout Chinese performance; excellent small models
DeepSeek-V3 / R1DeepSeek2024-12 / 2025-01671B (≈37B active)YesExtremely cheap to train; the open reasoning-model benchmark — see DeepSeek-R1 and Reasoning Models
Mistral 7B / Mixtral 8x7BMistral AI2023-09 / 2023-127B / 8x7B (MoE)YesEfficient and lightweight; Europe's open-source standard-bearer
text
Three rules of thumb for reading this table
1) Parameter count is no longer the only metric: MoE architectures routinely
   pair "huge total parameters" with "small active parameters" — judge
   capability and cost by active parameters, not total.
2) Closed-source models win on the polished end-to-end experience; open-source
   models win on control and cost: they can be deployed on-premises, keep data
   in-house, and be fine-tuned further.
3) Nobody stays number one: leaderboards reshuffle every couple of months.
   Anchor your selection to "your task," not to "the strongest."

Core Strengths and Core Limits ​

Core strengths ​

  • Generation and creation: drafting, rewriting, translating, summarizing, brainstorming.
  • Knowledge Q&A: covers factual knowledge in the training corpus (but only up to the training cutoff).
  • Reasoning and planning: solves math and logic problems via CoT; serves as the "brain" of an agent for task decomposition.
  • Code: generating, explaining, and debugging code — which spawned killer apps like GitHub Copilot and Code Intelligence.

Core limits ​

LimitHow it shows upTypical countermeasures
HallucinationConfidently invents nonexistent facts, fabricated citations and linksGround it with RAG, plus evaluation and human spot checks — see Retrieval-Augmented Generation and LLM Evaluation and Benchmarks
Limited context windowAttention compute grows quadratically with length; very long documents "won't stay in memory"Long-context models, layered summarization, RAG to feed only relevant chunks
High inference costEvery generated token runs a forward pass; more tokens = slower and pricierQuantization, speculative decoding, caching — see Inference Optimization and Quantization
Stale knowledgeKnows nothing past the training cutoffPlug into search engines or live data feeds — see Perplexity and AI Search
Safety risksJailbreaks, bias, privacy leaks, manipulation into harmful outputAlignment, guardrails, content filtering — see AI Safety and Governance

The costliest lesson

Hallucination does not disappear as models grow — it gets harder to spot because the model sounds more confident. Anywhere a wrong answer carries a real cost — healthcare, finance, legal, code headed to production — never let an LLM run naked; layer on retrieval, verification, and human review. For the related engineering traps, see Common Pitfalls and Anti-Patterns.

The Application Landscape: Prompt / RAG / Agent ​

Out of the box, an LLM is just a "probability generation machine." Turning it into business value means wrapping it in different "working modes." The three dominant paradigms today:

DimensionDirect call (Prompt)Bolt-on knowledge (RAG)Autonomous execution (Agent)
Core ideaWrite the question clearly; let the model answer directlyRetrieve relevant material first, then answerLet the model plan and call tools to finish the job
CostLowestMedium (+ vector DB and retrieval pipeline)High (multi-turn calls + tools + error handling)
Factual accuracyPoor (highest hallucination risk)Good (answers can cite sources)Medium (depends on what the tools actually return)
Knowledge freshnessPoor (frozen at the training cutoff)Good (the index can be updated in real time)Good (can query/crawl live)
Private dataNoYes (knowledge-base index)Yes
Task complexityOne-shot Q&A / generationMostly single-turn Q&AMulti-step planning, decomposition, execution
Typical stackAPI + promptsVector DB + retrieval + rerankingFunction calling + browser/code execution

Where Each Approach Fits ​

  • Direct call (Prompt): best for "one-shot, open-ended, no hard factual requirement" tasks — copywriting, brainstorming, polishing, translation. The fastest to get started with; techniques in Prompt Engineering and The Prompt Playbook.
  • RAG (Retrieval-Augmented Generation): best for tasks where "answers have correct sources and provenance matters" — customer-service Q&A, enterprise knowledge bases, live information. It bolts retrieval onto the LLM and is the de facto standard for enterprise deployments today. Build one from scratch in Build a RAG App from Scratch; the underlying semantic-retrieval principles in Vector Databases and Semantic Search.
  • Agent: best for "multi-step, hands-on" tasks — research a topic and write a report, schedule automatically, operate software. The LLM plans; tools execute. See AI Agents, Build an Agent from Scratch, and Manus and Agent Apps.

One-line selection test

Ask three questions: do the answers need verifiable sources? Does the knowledge need to be current? Does the task need multiple steps? All no → Prompt; yes to sources/freshness → RAG; yes to multi-step → Agent. Most real products are "RAG as the foundation, with agents at the critical steps."

Common Misconceptions: An LLM Is Not a Database, nor a Search Engine ​

Treating an LLM like a database to query or a search engine to search is the most common pair of rookie mistakes. The three are fundamentally different:

DimensionLLMDatabaseSearch engineRAG (LLM + retrieval)
Underlying operationProbabilistic generationExact queriesKeyword/relevance matchingRetrieval + generation
What it returnsNatural-language textStructured recordsLists of linksAnswers with citations
On a known factMay fabricateAlways accuratePoints to sourcesAs accurate and traceable as it can be
Knowledge updatesAt training timeWritten in immediatelyCrawled continuouslyJust update the index
  • An LLM is not a database: a database stores "definite facts," so whatever you query out is true. An LLM stores "statistical regularities" — it only guarantees "sounds right," not "is right." Business ground truth belongs in a database; the LLM's job is to translate it into natural language.
  • An LLM is not a search engine: a search engine returns sources and you make the judgment call; an LLM returns conclusions and makes the judgment for you — and when it is wrong, you may not even notice where. "Search-engine freshness + LLM expressiveness" is exactly RAG's territory, and it also drives hybrid directions like Recommender Systems in the LLM Era and Knowledge Graphs and Knowledge Injection.

How LLMs Relate to Multimodal and Diffusion Models ​

The LLM's "text-only" nature is a historical limitation, not an essential one:

  • Multimodal: extend next-token prediction from text to "tokens for images, audio, and video," and the same Transformer can look at pictures, listen, and talk about what it sees — that is how natively multimodal models like GPT-4o and Gemini took shape. See Multimodal Models.
  • Diffusion models: image/video generation is dominated by Diffusion Models and Generative AI. They complement rather than replace LLMs — language handles "understanding and planning," diffusion handles "pixel rendering." Video models like Sora merge the two into a single framework; see Video Generation and Image Generation.

The big-picture view

From 100,000 feet up: the LLM is the thinking brain; RAG and databases are the bookshelves; search engines and tools are the hands and feet; multimodal and diffusion are the senses and the paintbrush. For the overall architecture that assembles them, see Anatomy of the Overall Architecture.

Further Reading ​

References ​