Skip to content

What Is a Large Language Model

At a glance Large Language Models (LLMs) are neural network language models pre-trained on massive text corpora, with "predict the next word" as their fundamental task, typically having billions of parameters or more. Their capabilities emerge from thecombination of scale, data, and alignment.

What Is a Large Language Model ​

One-sentence definition: A Large Language Model (LLM) is a neural network language model pre-trained on massive text, with "predict the next word" as its core task and typically having billions of parameters or more — it may look like a "chatbot," but at its core, it's an extremely optimized next-word predictor. Its capabilities don't come from design; they emerge from the combination of scale (parameters and data volume), data (massive and diverse text), and alignment (making outputs conform to human preferences).

Remember one paradox first

LLMs have an incredibly simple training task — "guess the next word" — yet they exhibit incredibly rich capabilities — writing code, reasoning, following instructions, handling multi-turn conversations. Simple task × massive scale = complex capabilities. This is the most counterintuitive, yet most fundamental, fact about large language models — and the recurring theme this book returns to.

I. Three Layers of Definition ​

The one-sentence definition only answers "what is it." To truly understand LLMs, you need to look at three layers: what it does (practical layer), how it works (mechanism layer), and what it can achieve (capability layer).

1. Practical Layer: A "text in, text out" interface ​

From a user's perspective, an LLM is a function that takes text and returns text. You input an instruction (prompt), and it outputs a completion:

text
Input (prompt):     Explain attention mechanism in one sentence.
Output (completion): Attention mechanism lets the model dynamically focus on relevant parts of the input sequence when predicting each word.

Input:  def fibonacci(n):        # Ask the model to complete the code
Output:     if n <= 1: return n
                return fibonacci(n-1) + fibonacci(n-2)

The practical layer determines an LLM's product form: Q&A, writing, coding assistant, translation, summarization, conversation — all of these are the same thing at the interface level: a text-to-text mapping. This is also why it's called a "foundation model": one model, paired with different prompts and tools, can produce countless applications.

2. Mechanism Layer: An autoregressive model that "predicts the next word" ​

Behind the interface, an LLM has only one core operation: given preceding text, predict the probability distribution of the next token. A token is the model's smallest text processing unit (it can be a word, part of a word, or a character) — see Tokenization and Vocabularies.

In probabilistic terms, a language model estimates the probability of a piece of text:

text
P(w₁, w₂, …, wₙ) = P(w₁) · P(w₂|w₁) · P(w₃|w₁,w₂) · … · P(wₙ|w₁,…,wₙ₋₁)

That is, it splits the probability of a full sequence into successive conditional probabilities, each predicting just the next word. During training, the model reads massive amounts of text, repeatedly doing the same thing: see preceding text → predict the next word → compare with the actual next word → update parameters to predict more accurately. This goal of "maximizing the probability of real text" is the language modeling paradigm. It's simple enough to state in one sentence, but it only truly scaled with the Transformer architecture — because Transformer made "parallel processing of the full sequence + capturing long-range dependencies" feasible for the first time.

Why does "predicting the next word" teach knowledge?

This is the biggest question for beginners. The intuition is: to predict the next word accurately, the model must first understand the preceding text. To predict "After Newton discovered __, he developed calculus," the model must "know" about Newton, gravity, calculus, and the relationships between them. When the training text covers most of humanity's written expression of knowledge, the optimal way to compress that text is to encode those "knowledge" representations into the model's parameters. Predicting the next word is a surrogate task — on the surface, you're learning language; in reality, you're learning the structure of world knowledge. See Language Modeling for details.

3. Capability Layer: Emergent task-solving abilities ​

The mechanism layer only promises "predict the next word," but the capability layer shows what the model can ultimately do — far beyond "fill in the blank":

CapabilityManifestationUnderlying mechanism
Language understandingComprehend instructions, summarize, classify, sentiment analysisLanguage and knowledge representations accumulated during pretraining
Text generationContinuation, creative writing, translation, rewritingAutoregressive sampling (see Inference Fundamentals)
In-context learning"Learn on the fly" after a few examplesPattern matching within the context window, without parameter changes
Instruction followingExecute requests like "write me…," "use JSON format…"Post-training alignment (SFT + RLHF, see Alignment)
Reasoning (emergent)Math, logic, code debugging — emerges as scale growsEmergence driven by scaling laws (see Scaling Laws)

"In-context learning" and "instruction following" are particularly worth noting: neither of these capabilities was directly required by the pretraining task. They emerge from scale amplification and are strengthened by alignment training. This is what distinguishes large models from traditional ones — the capability list cannot be fully anticipated at design time.

II. Inheritance: From Statistical Language Models to LLMs ​

LLMs didn't appear out of nowhere. They are the direct descendants of three generations of language models:

GenerationRepresentativeCore ideaLimitations
Statistical language modelN-gram (1950s–2000s)Estimate the next word using co-occurrence frequency of the previous n-1 wordsData sparsity, short window, cannot generalize to unseen word sequences
Neural language modelBengio 2003 neural language model → RNN/LSTMUse neural networks to map words to vectors, model arbitrary-length contextSlow sequential processing, still difficult with long-range dependencies, limited scale
Large language modelTransformer pre-trained large models (2018–present)Massive data + large-scale parallel architecture + parameter scale growthHigh training cost, hallucination, alignment difficulty

There are two key turning points: Bengio's neural language model in 2003 (using neural networks instead of counting), and the 2017 Transformer and 2018 GPT/BERT pretraining paradigm (pretrain on massive corpora first, then adapt to downstream tasks). By 2020, GPT-3 proved that "when a model is large enough, tasks can be solved via prompting without fine-tuning," and "large language model" officially became an independent concept. See Brief History for the full timeline.

1. A Continuing Thread: One Example ​

See how three generations of models approach the same goal. The task is to predict the next word of "The cat sat on the ____":

ModelHow it "guesses"Result
N-gram (statistical)Counts what word most commonly follows "on the" in the corpusOften predicts objects, but cannot use more distant context, and unseen combinations get zero probability
Neural language model (Bengio/RNN)Maps words to vectors; similar semantics share statisticsCan generalize to unseen combinations, but limited by sequence length
LLM (Transformer)Attends to the full context, combining context and knowledgeCan infer reasonable words like "mat/rug" from "cat, sat, on," and also consider worldworld knowledge

The three generations share the same objective function (maximizing real text probability). The differences are entirely in representational capacity (from counting tables to continuous vectors to deep representations) and modeling scope (from n words to the full context). That's the true meaning of "inheritance": not a different problem, but different tools for the same problem.

III. Capability Inventory ​

Separating "usable" from "reliable" capabilities, here's an LLM capability map:

Capability domainRepresentative capabilitiesTypical behaviorDifficulty
LanguageUnderstanding, generation, rewriting, translation, summarizationStable and usable, works in both Chinese and English★
KnowledgeFact memorization, commonsense reasoningPowerful but outdated and error-prone★★
ReasoningMath, logic, codeImproving with scale and reasoning-time scaling★★★
InteractionInstruction following, multi-turn dialogue, role-playingDepends on alignment quality★★
Tool useFunction calling, search/calculator/browser accessDepends on external orchestration (see Agents with LLMs)★★★
MultimodalImage understanding, audio processing, image generationNative multimodal models are becoming mainstream★★★
text
The "emergence" curve of capabilities (illustrative, not real data):

Capability level
  │                                    ● Reasoning (emergent, jumps after threshold)
  │                            ●
  │                    ●   Language understanding (smooth rise)
  │            ●
  │    ●
  └──────────────────────────────────────→ Model scale

Capabilities don't come for free

"Stronger capability" usually means "more parameters, more training data, higher cost." Strong capability and high cost often go hand-in-hand. There's also an evaluation gap between "high scores" on benchmarks and "usefulness" in production — see Evaluation and Benchmarks.

1. Practical Implications of Capability Tiers ​

If you further slice capabilities by "maturity," it directly impacts engineering decisions:

TierRepresentative capabilitiesProduction readinessPractical implications
MatureTranslation, summarization, rewriting, structured extractionHigh, near ready-to-deployReady to use; main work is prompting and evaluation
DevelopingCode generation, math reasoning, multi-step planningModerate, needs guardrails and verificationUse with testing, sandboxes, and human review
ExperimentalComplex long-horizon tasks, autonomous agentsLow, insufficient reliabilityStart small, gradually expand (see Agents with LLMs)

The same model has different capabilities at different maturity levels — it may be reliable at translation but unreliable at autonomous planning. Mature practice is: pick tiers for tasks, pair tiers with guardrails, rather than entrusting everything to a "smart model." This "capability tier thinking" runs through the "application" phase of Overall Architecture Anatomy.

IV. Key Components ​

An LLM is the product of five elements, all essential:

ElementDescriptionExample scale (2025 mainstream reference)Related pages
ParametersNumber of learnable parameters, determines capacityBillions to hundreds of billions (e.g., Llama-3-405B); MoE models have larger total but fewer activated parametersScaling Laws, MoE
Training dataPretraining corpus, determines knowledge scope and qualityTrillions of tokens (mixed multilingual, code, math, books)Pretraining
ArchitectureMostly Decoder-only TransformerAutoregressive + causal masking + positional encodingTransformer Architecture
AlignmentMakes outputs conform to human preferences (helpful, honest, safe)SFT + RLHF / DPO and other post-training methodsAlignment
Context windowHow much text the model can "see" at onceThousands to hundreds of thousands of tokensContext and Long Contexts

Memory aid

Large model = large parameters × large data × good architecture × good alignment × long context. When asked "why are large models 'large'?" in an interview, list the five elements, then land on "thecombination of scale, data, and alignment."

V. A Minimal Example: Run a Next-Token Generation by Hand ​

Ground the abstract in code. Use Hugging Face transformers to load an open-source small model (GPT-2 small, ~124M parameters; the full GPT-2 is 1.5B parameters) and do one "predict the next word":

python
# Prerequisite: pip install transformers  (model weights download automatically on first run)
from transformers import AutoModelForCausalLM, AutoTokenizer

# ① Load a small open-source model (GPT-2 small, 124M parameters, enough to demo "next-word prediction")
model_name = "gpt2"
model = AutoModelForCausalLM.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)

# ② Input preceding text (use a Chinese base model for Chinese; here we use English for easy verification)
prompt = "The capital of France is"
inputs = tokenizer(prompt, return_tensors="pt")

# ③ The model outputs probabilities for each candidate token; take the highest as "predicted next word"
outputs = model.generate(**inputs, max_new_tokens=1)
print(tokenizer.decode(outputs[0]))
# Expected output: ... France is Paris    ← the model "predicted" Paris

Try changing the prompt to "2 + 2 =" or "Once upon a time", and observe that the model only moves forward one token at a time, with output being generated word by word — this is autoregression. To understand the mechanisms behind this process (sampling, temperature, etc.), see Inference Fundamentals; to train a smaller model from scratch, see Building a Large Model from Scratch.

Hands-on suggestion

Repeat the code above with the smallest Chinese model your computer can run (e.g., Qwen2.5-0.5B), change max_new_tokens to 50, and observe a full continuation. Ten minutes of hands-on experimentation beats ten readings — you'll immediately grasp the relationship between "predicting the next word" and "writing an entire paragraph."

1. From Example to Production: Three Leaps ​

The demo above is only 30 lines of code, and production systems are three orders of magnitude away. Understanding this gap helps you calibrate the distance between "example" and "engineering":

LayerMinimal example (above)Production systemSource of gap
ModelSmall model (100M parameters)Large model (billions to hundreds of billions of parameters, or API)Scale and capability (see Scaling Laws)
InteractionSingle completion, no memoryMulti-turn dialogue, tool calling, system promptsContext management (see Context and Long Contexts)
DeliveryCLI printoutLow-latency, high-concurrency, evaluated and monitored serviceDeployment engineering (see Deployment and Servicing)

The value of an example is personally confirming the mechanism; the foundation of production is systematic engineering. Don't underestimate production because "the example is so simple," and don't skip the example because "production is so complex."

VI. Why Large Models Work: Three Empirical Benchmarks ​

"Scale, data, and alignmentcombination" is not a slogan — it's an empirical chain that can be repeatedly observed. The three benchmarks successively demonstrate the power of scale, the breakthrough of alignment, and the limits of capability:

BenchmarkEventSignificance
GPT-3 Few-shot (2020)GPT-3 with 175B parameters, without fine-tuning, nearly matched or exceeded the best fine-tuned results on many tasks using only a few examples in prompts (few-shot)Proved that "scale itself brings capability," and task adaptation can rely on prompting rather than training; see GPT Series
ChatGPT's Breakout (Nov 2022)A conversational product built on InstructGPT's RLHF technology hit 1 million users in 5 days and 100 million in two monthsProved that "capability + alignment + conversational interaction" can cross professional boundaries and let ordinary people use it directly; see ChatGPT and Conversational Models
GPT-4 on Exams (2023)GPT-4 placed in the top 10% of humans on the Uniform Bar Exam and several other testsProved that models can reach expert-level performance on "tasks requiring reasoning and integrated knowledge," and also turned "evaluating LLMs" itself into a discipline (see Evaluation and Benchmarks)

These three steps correspond to the three threads this book repeatedly emphasizes: scale (GPT-3) → alignment (ChatGPT) → capability boundaries and evaluation (GPT-4).

1. After the Three Benchmarks: New Dimensions in 2024–2025 ​

Advances after 2024 can be seen as adding a dimension to each of these three benchmarks:

DimensionNew contentRepresentativeLink to
Scale dimensionInference-time scaling (test-time scaling): counting "thinking time" as part of scaleo1 series, DeepSeek-R1Frontier Progress
Alignment dimensionFrom text alignment to multimodal and tool-behavior alignmentGPT-4o, GeminiMultimodal LLMs
Evaluation dimensionFrom benchmark scores to long-context, Agent tasks, and safety evaluationLongBench, Agent benchmarksEvaluation and Benchmarks

These new dimensions don't overturn the "scale, data, alignment" framework; they keep updating the content within each block of that framework — which also explains why this handbook continuously follows up in Frontier Progress and Mainstream Model Profiles.

VII. Trade-offs and Boundaries ​

Beyond the LLM capability map, there's also an "unreliability map." Being honest about boundaries is the only way to use LLMs well:

BoundaryManifestationMitigation directionRelated pages
HallucinationConfidently fabricating facts, citing non-existent papersRetrieval-augmented generation (RAG), factuality evaluation, prompt constraintsHallucination, RAG
CostCompute, memory, electricity, API fees for training and inferenceQuantization, distillation, MoE sparsification, on-demand model selectionDeployment and Servicing
Evaluation difficulty"Right or wrong" of generative outputs is hard to judge automatically; leaderboards can be gamedMulti-dimensional evaluation, LLM-as-a-judge, human reviewEvaluation and Benchmarks
Knowledge stalenessPretraining data has a cutoff date; the model can't know what happened afterRAG, real-time search tools, periodic model updatesContext and Long Contexts
Safety and biasHarmful outputs, privacy, bias amplification, vulnerability to prompt injectionAlignment training, guardrails, content moderationSafety and Risks

Two common misjudgments

  • "LLMs lie" is a misunderstanding. The model has no intent to "lie"; it's just optimizing "the probability of the next word." Hallucination is a statistical machine's inappropriate confidence under uncertainty, not a personality flaw.
  • "LLMs know everything" is an even bigger misunderstanding. It's a "language distribution compressor," skilled at language-formatted knowledge but weak at tasks requiring real-time world state, precise computation, and real causality. Using it as a database, a calculator, or a source of truth will all lead to disappointment.

1. Three Principles for Using LLMs ​

Facing the above boundaries, mature engineering usage can be distilled into three principles:

  1. Give the model facts instead of making it recall them: For knowledge-intensive tasks, prefer RAG or tool calling, letting the model "answer based on given materials." See Hallucination and RAG.
  2. Design verification for critical decisions: For outputs requiring precision (code, numbers, structured data), fall back on rule validation, unit tests, or human review — don't trust the model's confidence.
  3. Treat evaluation as part of the product: Define "what counts as a correct answer" before launch, continuously sample for regression after launch, and make every model change measurable. See Evaluation and Benchmarks.

One-sentence summary

An LLM is a general-purpose capability engine with language as its interface, shaped by scale and alignment, to be used within acceptable boundaries. Treating it as a "smart colleague" rather than an "omniscient database" is the attitude this page most wants to convey.

VIII. How This Book Unfolds ​

This page is the master entry point for all concepts. We recommend continuing with Learning Paths: first read Brief History to build a timeline, use Overall Architecture Anatomy to build a site-wide map, then go deeper along the main thread of Language Modeling → Transformer Architecture → Pretraining → Alignment. Whenever you encounter a term, go back to the Glossary.

Further Reading ​

  • Language Modeling — Full expansion of the "mechanism layer" from this page: next-word prediction paradigm, perplexity, and cross-entropy
  • Transformer Architecture — The architectural revolution that made "massive text pretraining" feasible: QKV and multi-head attention
  • Scaling Laws — A quantitative answer to "why bigger models are better," with discussion of emergent capabilities
  • Alignment — How "helpfulness, honesty, and safety" enter the model through RLHF/DPO
  • Brief History — Three paradigm shifts from N-gram to ChatGPT
  • Inference Fundamentals — Autoregressive sampling, temperature and top-p, understanding how models "speak"

References ​