Appearance
Large Language Models (LLMs)
Definition in One Sentence
A large language model (LLM) is a language model built on the Transformer architecture, pretrained on massive amounts of text with "predict the next token" as its learning objective, and usually packing billions of parameters or more.
It is both a "language" model and a "world" model: to predict the next word, the model is forced to distill grammar rules, factual knowledge, logical relations, even persona and tone, out of billions of web pages, books, and code. The Transformer paper landed in 2017, GPT and BERT followed in 2018, and ChatGPT detonated the mass market in late 2022. LLMs are now the core engine of the generative AI wave — most of the hot topics in What Are AI's Hot Concepts revolve around them.
Who uses LLMs? Practically everyone, just in different ways: ordinary people ask questions of chat products; developers build applications through APIs; researchers keep making them bigger, smaller, and better understood; enterprises use them to reshape customer service, writing, code, search, and nearly every other business function. It is no exaggeration to say that the LLM is the single most significant AI breakthrough of the past decade. To see where it sits on the wider technical map, start with The Concept Landscape and A Brief History of AI.
The Core Mechanism: Autoregressive Language Modeling
Everything an LLM can do boils down to one operation: given a piece of text, predict what the most likely next token is.
From Tokens to Conditional Probability
Raw text is first split into tokens — words, subwords, or characters. The model then frames "predict the next token" as a conditional probability:
text
P(xₜ | x₁, x₂, …, xₜ₋₁)Given the first t-1 tokens, predict a probability distribution over the t-th token. The joint probability of the whole text is the product of every token's conditional probability:
text
P(x₁, x₂, …, xₙ) = ∏ₜ P(xₜ | x₁, …, xₜ₋₁)Training maximizes this log-likelihood (equivalently, minimizes the cross-entropy loss):
text
L = -Σₜ log P(xₜ | x₁, …, xₜ₋₁)Training and Inference: Two Sides of the Same Coin
The elegance of this mechanism is that at training time the model learns nothing but "continue the text," and at inference time that is exactly all it does — both share the same code path. During training, real text serves as the answer key for step-by-step prediction (teacher forcing); during inference, each token the model just generated becomes part of the input for the next step, and the whole output is produced autoregressively, one token at a time:
python
# Autoregressive generation: one token at a time
def generate(model, prompt, max_new_tokens=100):
tokens = tokenize(prompt)
for _ in range(max_new_tokens):
logits = model(tokens) # the model outputs a distribution over the next token
next_token = sample(logits[-1]) # sample from the distribution (temperature/top-p controllable)
tokens.append(next_token)
if next_token == EOS: # stop at the end-of-sequence token
break
return detokenize(tokens)Key takeaway
An LLM is not a system that "looks up answers by keyword." It is a generator of probability distributions over a vocabulary. Every generation is a sample — the same question can get different answers on different runs. That is not a bug; it is an inherent property of a probabilistic model. Keep this in mind as the starting point for understanding every capability and flaw that follows, hallucination above all.
"Predict the next token" sounds simple, but for predictions to be accurate, the model must internalize vast amounts of knowledge. The Transformer's self-attention lets every token "see" every other token in the context, which is what makes modeling long-range dependencies possible — the physical foundation of LLMs. Details in Transformers and Attention.
The Three-Stage Training Paradigm: Pretraining → Post-Training
An LLM's abilities are not acquired in a single pass; they are cultivated in stages. Modern LLMs follow a "pretraining + post-training" paradigm, and post-training itself splits into supervised fine-tuning and alignment:
text
Massive text corpora (trillions of tokens: web pages, books, code)
│
▼
┌────────────────────────────────┐
│ (1) Pretraining │
│ Goal: predict the next token │
│ Output: a base model that can │
│ "continue text" │
└────────────────────────────────┘
│
▼
┌────────────────────────────────┐
│ (2) Supervised Fine-Tuning │
│ (SFT / instruction tuning)│
│ Data: human-written │
│ instruction-answer pairs │
│ Output: organizes knowledge │
│ into proper "answers" │
└────────────────────────────────┘
│
▼
┌────────────────────────────────┐
│ (3) Alignment (RLHF / DPO) │
│ Data: human preference │
│ rankings over responses │
│ Output: a model that is safe, │
│ honest, and helpful │
└────────────────────────────────┘
│
▼
Deploy for inference- Pretraining: next-token prediction over massive unlabeled text. It takes weeks and tens of millions of dollars and produces a "base model." The base model already has language ability and a huge store of knowledge, but it "only continues text — it doesn't answer questions properly."
- Supervised Fine-Tuning (SFT): continued training on tens of thousands of "instruction → desired answer" pairs, teaching the model to answer whatever the user asks. This is the step that turns a text-continuing infant into an assistant that answers questions; the mechanics are covered in Fine-Tuning and PEFT (LoRA).
- Alignment: trained on human preference data (rankings over multiple responses) via RLHF or DPO, making the model readier to refuse harmful requests, more honest about what it doesn't know, and less sycophantic. Methods and costs in Alignment: RLHF and DPO.
One sentence to remember
Pretraining sets the ceiling on capability; post-training sets the behavioral style. Post-training cannot conjure knowledge the base model never learned — it can only shape existing abilities into what users want. Many "the model got dumber" complaints trace back to alignment suppressing the model's expressiveness too aggressively.
If you plan to reproduce this pipeline yourself (even at small scale), follow the hands-on route in Fine-Tune Your Own LLM.
Emergent Abilities: When Scale Changes Quality
Once model scale crosses a certain threshold (usually somewhere past tens of billions of parameters), abilities that were never explicitly required by the training objective suddenly appear. These are emergent abilities:
| Ability | One-line description | How you trigger it |
|---|---|---|
| In-context learning | Learns a new task from a few examples in the prompt, with no parameter updates | Provide 1–5 input-output examples in the prompt |
| Instruction following | Understands and executes natural-language instructions | Describe the task in natural language |
| Chain-of-thought (CoT) reasoning | "Think step by step, then answer" markedly improves math/logic accuracy | Ask the model to "think step by step" |
| Few-shot | Handles never-before-seen tasks from a handful of examples | Embed examples in the prompt |
These abilities are nothing like the classical "train one model per task" playbook, and they are precisely the watershed separating LLMs from every NLP model that came before. The research evidence comes from Wei et al. (2022), Emergent Abilities of Large Language Models — see Core Papers, Close Up.
How Emergence Relates to Scaling Laws
Scaling laws describe a predictable power-law relationship between "bigger model + more data + more compute" and falling loss: Kaplan et al. (2020) found that loss falls as a power law in parameters, data, and compute; Chinchilla (Hoffmann et al., 2022) then derived the "compute-optimal" recipe — when model parameters double, training tokens should double as well, which is where today's rule of thumb of "roughly 20 tokens per parameter" for training large models comes from.
The relationship between the two fits in one line:
The one-line verdict
Scaling laws deliver predictable, smooth improvement; emergence delivers unpredictable, discontinuous jumps. Scale pushes loss steadily downward, while certain abilities suddenly "unlock" at particular points on the loss curve — we cannot precisely predict at what scale they will appear, only stumble into them experimentally. That is the deep motive behind frontier teams still racing to pile on scale; see Frontier Progress.
This also explains why the 2020s "arms race" is so fierce: whoever reaches the next threshold scale first unlocks the next batch of abilities first.
Architecture Families: From BERT and GPT to T5
Depending on how the Transformer is used, language models fall into three families:
| Family | Representative models | Attention | Pretraining objective | Good at | Where they stand today |
|---|---|---|---|---|---|
| Encoder-only | BERT, RoBERTa | Bidirectional | Masked language modeling (MLM) | Understanding, representations, classification, retrieval | The workhorse for semantic vectors / embeddings |
| Decoder-only | GPT family, Llama, Qwen, DeepSeek | Unidirectional (causal) | Next-token prediction | Generation, dialogue, reasoning | The absolute mainstream |
| Encoder-Decoder | T5, BART | Bidirectional encoding + autoregressive decoding | Span infilling | Translation, summarization, other transduction tasks | Still valuable for specific tasks |
Why Decoder-only Won in the End
In 2018, BERT (a bidirectional encoder) crushed GPT-1 on understanding tasks, and the industry briefly believed "encoders are the way." From 2023 on, virtually every flagship model has been Decoder-only. Three reasons:
- Training and inference objectives match: generation is the only way an LLM produces output, and the causal (unidirectional) next-token training objective is exactly isomorphic to autoregressive inference — no "encoder understands, decoder generates" objective mismatch.
- Engineering efficiency: causal attention is a natural fit for inference acceleration (KV cache reuse), and the Decoder-only architecture is simpler and easier to scale to enormous sizes.
- Task unification: translation, summarization, Q&A, coding — all of it can be expressed as text-in, text-out sequence-to-sequence tasks, which a single Decoder-only model handles without any per-task head designs.
Fun fact
T5's own conclusion ran opposite to today's consensus: it found encoder-decoder slightly better than decoder-only on transduction tasks. History showed that decoder-only models overtook it on the strength of scale and ecosystem (OpenAI's GPT line simply kept getting bigger). The final winner of an architecture debate is usually decided by engineering scalability and ecosystem, not single-benchmark accuracy.
For the internals — attention, positional encodings, KV cache, and more — see Transformers and Attention.
Key Models Compared
As of mid-2025, the global LLM landscape is roughly "three big US closed-source players plus a four-way open-source contest." The table below exists to build intuition — models update monthly; for the latest leaderboards, check Models and Leaderboards Quick Reference.
| Model | Maker | Released | Parameters | Open source | Notes |
|---|---|---|---|---|---|
| GPT-4o | OpenAI | 2024-05 | Undisclosed (believed MoE) | No | Natively multimodal, low latency; ChatGPT's mainstay — see ChatGPT and Conversational AI |
| Claude 3.5 / 3.7 Sonnet | Anthropic | 2024-06 / 2025-02 | Undisclosed | No | Long context, strong code and long-form writing, standout safety alignment |
| Gemini 1.5 Pro / 2.5 Pro | 2024-02 / 2025-03 | Undisclosed | No | Million-token-scale context window, deep multimodal integration | |
| Llama 3.1 | Meta | 2024-07 | 8B / 70B / 405B | Yes | The open-weights flagship; largest ecosystem |
| Qwen2.5 | Alibaba | 2024-09 | 0.5B–72B | Yes | Multilingual with standout Chinese performance; excellent small models |
| DeepSeek-V3 / R1 | DeepSeek | 2024-12 / 2025-01 | 671B (≈37B active) | Yes | Extremely cheap to train; the open reasoning-model benchmark — see DeepSeek-R1 and Reasoning Models |
| Mistral 7B / Mixtral 8x7B | Mistral AI | 2023-09 / 2023-12 | 7B / 8x7B (MoE) | Yes | Efficient and lightweight; Europe's open-source standard-bearer |
text
Three rules of thumb for reading this table
1) Parameter count is no longer the only metric: MoE architectures routinely
pair "huge total parameters" with "small active parameters" — judge
capability and cost by active parameters, not total.
2) Closed-source models win on the polished end-to-end experience; open-source
models win on control and cost: they can be deployed on-premises, keep data
in-house, and be fine-tuned further.
3) Nobody stays number one: leaderboards reshuffle every couple of months.
Anchor your selection to "your task," not to "the strongest."Core Strengths and Core Limits
Core strengths
- Generation and creation: drafting, rewriting, translating, summarizing, brainstorming.
- Knowledge Q&A: covers factual knowledge in the training corpus (but only up to the training cutoff).
- Reasoning and planning: solves math and logic problems via CoT; serves as the "brain" of an agent for task decomposition.
- Code: generating, explaining, and debugging code — which spawned killer apps like GitHub Copilot and Code Intelligence.
Core limits
| Limit | How it shows up | Typical countermeasures |
|---|---|---|
| Hallucination | Confidently invents nonexistent facts, fabricated citations and links | Ground it with RAG, plus evaluation and human spot checks — see Retrieval-Augmented Generation and LLM Evaluation and Benchmarks |
| Limited context window | Attention compute grows quadratically with length; very long documents "won't stay in memory" | Long-context models, layered summarization, RAG to feed only relevant chunks |
| High inference cost | Every generated token runs a forward pass; more tokens = slower and pricier | Quantization, speculative decoding, caching — see Inference Optimization and Quantization |
| Stale knowledge | Knows nothing past the training cutoff | Plug into search engines or live data feeds — see Perplexity and AI Search |
| Safety risks | Jailbreaks, bias, privacy leaks, manipulation into harmful output | Alignment, guardrails, content filtering — see AI Safety and Governance |
The costliest lesson
Hallucination does not disappear as models grow — it gets harder to spot because the model sounds more confident. Anywhere a wrong answer carries a real cost — healthcare, finance, legal, code headed to production — never let an LLM run naked; layer on retrieval, verification, and human review. For the related engineering traps, see Common Pitfalls and Anti-Patterns.
The Application Landscape: Prompt / RAG / Agent
Out of the box, an LLM is just a "probability generation machine." Turning it into business value means wrapping it in different "working modes." The three dominant paradigms today:
| Dimension | Direct call (Prompt) | Bolt-on knowledge (RAG) | Autonomous execution (Agent) |
|---|---|---|---|
| Core idea | Write the question clearly; let the model answer directly | Retrieve relevant material first, then answer | Let the model plan and call tools to finish the job |
| Cost | Lowest | Medium (+ vector DB and retrieval pipeline) | High (multi-turn calls + tools + error handling) |
| Factual accuracy | Poor (highest hallucination risk) | Good (answers can cite sources) | Medium (depends on what the tools actually return) |
| Knowledge freshness | Poor (frozen at the training cutoff) | Good (the index can be updated in real time) | Good (can query/crawl live) |
| Private data | No | Yes (knowledge-base index) | Yes |
| Task complexity | One-shot Q&A / generation | Mostly single-turn Q&A | Multi-step planning, decomposition, execution |
| Typical stack | API + prompts | Vector DB + retrieval + reranking | Function calling + browser/code execution |
Where Each Approach Fits
- Direct call (Prompt): best for "one-shot, open-ended, no hard factual requirement" tasks — copywriting, brainstorming, polishing, translation. The fastest to get started with; techniques in Prompt Engineering and The Prompt Playbook.
- RAG (Retrieval-Augmented Generation): best for tasks where "answers have correct sources and provenance matters" — customer-service Q&A, enterprise knowledge bases, live information. It bolts retrieval onto the LLM and is the de facto standard for enterprise deployments today. Build one from scratch in Build a RAG App from Scratch; the underlying semantic-retrieval principles in Vector Databases and Semantic Search.
- Agent: best for "multi-step, hands-on" tasks — research a topic and write a report, schedule automatically, operate software. The LLM plans; tools execute. See AI Agents, Build an Agent from Scratch, and Manus and Agent Apps.
One-line selection test
Ask three questions: do the answers need verifiable sources? Does the knowledge need to be current? Does the task need multiple steps? All no → Prompt; yes to sources/freshness → RAG; yes to multi-step → Agent. Most real products are "RAG as the foundation, with agents at the critical steps."
Common Misconceptions: An LLM Is Not a Database, nor a Search Engine
Treating an LLM like a database to query or a search engine to search is the most common pair of rookie mistakes. The three are fundamentally different:
| Dimension | LLM | Database | Search engine | RAG (LLM + retrieval) |
|---|---|---|---|---|
| Underlying operation | Probabilistic generation | Exact queries | Keyword/relevance matching | Retrieval + generation |
| What it returns | Natural-language text | Structured records | Lists of links | Answers with citations |
| On a known fact | May fabricate | Always accurate | Points to sources | As accurate and traceable as it can be |
| Knowledge updates | At training time | Written in immediately | Crawled continuously | Just update the index |
- An LLM is not a database: a database stores "definite facts," so whatever you query out is true. An LLM stores "statistical regularities" — it only guarantees "sounds right," not "is right." Business ground truth belongs in a database; the LLM's job is to translate it into natural language.
- An LLM is not a search engine: a search engine returns sources and you make the judgment call; an LLM returns conclusions and makes the judgment for you — and when it is wrong, you may not even notice where. "Search-engine freshness + LLM expressiveness" is exactly RAG's territory, and it also drives hybrid directions like Recommender Systems in the LLM Era and Knowledge Graphs and Knowledge Injection.
How LLMs Relate to Multimodal and Diffusion Models
The LLM's "text-only" nature is a historical limitation, not an essential one:
- Multimodal: extend next-token prediction from text to "tokens for images, audio, and video," and the same Transformer can look at pictures, listen, and talk about what it sees — that is how natively multimodal models like GPT-4o and Gemini took shape. See Multimodal Models.
- Diffusion models: image/video generation is dominated by Diffusion Models and Generative AI. They complement rather than replace LLMs — language handles "understanding and planning," diffusion handles "pixel rendering." Video models like Sora merge the two into a single framework; see Video Generation and Image Generation.
The big-picture view
From 100,000 feet up: the LLM is the thinking brain; RAG and databases are the bookshelves; search engines and tools are the hands and feet; multimodal and diffusion are the senses and the paintbrush. For the overall architecture that assembles them, see Anatomy of the Overall Architecture.
Further Reading
- What Are AI's Hot Concepts — where LLMs sit in the whole AI landscape
- AI vs ML vs DL vs GenAI vs Agent — untangling the concept boundaries
- A Brief History of AI — four decades from n-grams to LLMs
- Transformers and Attention — the physical foundation of LLMs
- Fine-Tuning and PEFT (LoRA) — turning a base model into a specialized one
- Alignment: RLHF and DPO — making models safe, honest, and helpful
- Prompt Engineering — changing behavior without training
- Retrieval-Augmented Generation (RAG) — the right way to give an LLM a knowledge base
- AI Agents — from "can talk" to "can work"
- LLM Evaluation and Benchmarks — how to know a model is actually good
- Inference Optimization and Quantization — driving down inference cost
- ChatGPT and Conversational AI — where LLMs hit the mass market
- DeepSeek-R1 and Reasoning Models — the open reasoning-model benchmark
- Models and Leaderboards Quick Reference — mandatory reading before any model selection
- Learning Paths — which road to take for a systematic LLM education
References
- Vaswani et al., Attention Is All You Need (NeurIPS, 2017) — the original Transformer paper; the starting point of every modern LLM
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers (2018) — the encoder-only representative
- Raffel et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5, 2019) — the encoder-decoder representative
- Brown et al., Language Models are Few-Shot Learners (GPT-3, 2020) — the foundational paper on few-shot and in-context learning
- Kaplan et al., Scaling Laws for Neural Language Models (2020) — the founding work on scaling laws
- Hoffmann et al., Training Compute-Optimal Large Language Models (Chinchilla, 2022) — the classic parameter/data ratio result
- Wei et al., Emergent Abilities of Large Language Models (2022) — the paper that introduced "emergent abilities"
- Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022) — the original chain-of-thought (CoT) paper
- OpenAI, GPT-4 Technical Report (2023) — the closed-source flagship's technical report
- Touvron et al., LLaMA: Open and Efficient Foundation Language Models (2023) — the start of the open large-model wave
- Touvron et al., Llama 2: Open Foundation and Fine-Tuned Chat Models (2023) — the complete open-source + alignment (RLHF) template
- DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025) — the open reasoning-model benchmark
- OpenAI Platform docs: Models overview — official model list and pricing
- Hugging Face Open LLM Leaderboard — continuously updated open-model rankings