Theme
LLMs and Adjacent Concepts
One-sentence positioning: Large language models are a specific application form within the deep learning family, trained using language modeling as their paradigm. They inherit the full legacy of NLP, spawn new systems like agents and RAG, but their boundaries are far smaller than "artificial general intelligence." This paper places LLMs in a conceptual coordinate system and clarifies one by one "who contains whom, who depends on whom, and who is mistaken for whom."
I. Overview Table
| Concept | In one sentence | Relationship to LLM | Common misconception |
|---|---|---|---|
| Natural Language Processing (NLP) | The research field of enabling machines to process human language | LLMs are the current dominant implementation method of NLP | Thinking "NLP = LLM," when NLP also includes tokenization, syntax, word embeddings, knowledge graphs, and more |
| Deep Learning (DL) | A family of techniques using multi-layer neural networks for representation learning | LLMs are one application form of deep learning on text | Thinking "deep learning = LLM," when DL also includes CNNs, RNNs, reinforcement learning agents, etc. |
| Statistical language model | Estimating text probability using statistical methods (n-gram, etc.) | LLMs' direct ancestor — idea inherited, implementation upgraded | Thinking the two are unrelated, when "predicting the next word" shares a single lineage |
| Foundation Model | A large model pre-trained on massive data, adaptable to many tasks (including non-text) | LLMs are the text branch of foundation models (there are also vision, multimodal, etc.) | Treating "foundation model" as synonymous with LLM |
| Agent | A system that uses an LLM as its brain, paired with tools and planning loops | LLMs are the core component/substrate of agents | Thinking "using an LLM equals building an agent," when agents need orchestration, tools, and memory |
| RAG system | LLM + external retrieval, using retrieved results to assist generation | RAG is an augmentation paradigm for LLMs | Thinking RAG is a different type of model, when it's actually a system built around an LLM |
| Artificial General Intelligence (AGI) | General intelligence that reaches human-level performance on all tasks | LLMs are one candidate route to AGI, far from AGI itself | Equating "ChatGPT is smart" with "AGI has been achieved" |
A diagnostic tool
When concepts get confused, ask three questions: Is it a model or a system? Is it a family or an individual? What problem is it trying to solve? LLM = "model, individual, language task." Agent = "system, composed of LLMs, task execution." DL = "family, LLMs belong to it." AGI = "goal, not a specific product." After asking these three questions, most confusion dissolves automatically.
II. LLMs vs. Traditional NLP
1. Paradigm differences
Traditional NLP is "task decomposition + feature engineering": break language understanding into sub-tasks like tokenization, POS tagging, syntax parsing, named entity recognition — each task modeled separately with separate labeled data. LLMs are "unified paradigm + few-shot adaptation": one model, one pretraining objective, with downstream tasks handled via prompting, few-shot examples, or even zero-shot.
| Dimension | Traditional NLP | LLM era |
|---|---|---|
| Task form | One model per task | One model solves almost all text tasks |
| Data dependency | Each task needs a large amount of labeled data | Pretraining uses unlabeled corpora; downstream needs little or no labeled data |
| Language unit | Word | Token (subword) |
| Context | Local features (window, syntax tree) | Long context + attention |
| Representative methods | CRF, SVM, LSTM+Attention, BERT fine-tuning | Generative pretraining + prompting/alignment |
| Upgrade method | Change model, add features | Add data, add parameters, change prompts |
Traditional NLP's achievements haven't been discarded — they've been absorbed into the "implicit knowledge" of LLMs. Most post-training LLMs have internalized capabilities like tokenization, syntax, and entity recognition. Independent traditional NLP work mainly persists in three scenarios: resource-scarce languages, tasks requiring precise structured output (table extraction, information retrieval ranking), and scenarios with hard interpretability requirements.
A concrete example illustrates the difference. Named entity recognition (NER) is a classic sequence labeling task in traditional NLP: using BIO tags to mark "John went to New York yesterday" as "John=person, New York=location," requiring dedicated labeled data and models. An LLM's approach is to simply ask: "Extract person names and locations from this sentence," and the model gives an answer in seconds, with zero-shot adaptation to new entity types. The tradeoff: traditional methods are stable, controllable, and cheap; LLMs are flexible, general-purpose, but occasionally drift. Production systems often pair them as "LLM baseline + rules/small models as fallback."
2. Why "LLM became the synonym for NLP in the 2020s"
One noteworthy phenomenon: after 2020, the submission topics at major NLP conferences (ACL, EMNLP, NAACL), industrial job positions, and academic hiring have been almost entirely dominated by LLMs. There are three reasons:
- Capability coverage: LLMs have matched or exceeded dedicated models on most language tasks, replacing the old task-by-task checklist with "one model eats all."
- Engineering unification: Teams only need to maintain one "data + pretraining + alignment + inference" pipeline, rather than maintaining separate labeling and model systems for each task (compare with the lifecycle in Overall Architecture Anatomy).
- Talent market: Job titles shifted from "NLP engineer" to "large model algorithm engineer." The skill tree moved from "feature engineering" to "prompting, fine-tuning, evaluation, RAG/Agent orchestration." See Careers & JD.
Be careful with wording
"Saying 'LLM replaced NLP' is an exaggeration. More accurately: LLM replaced most of NLP's application-layer implementations, but NLP as a discipline studying the nature of language (language structure, semantics, pragmatics, multilingual, low-resource) still exists independently — only its outputs now mostly appear in the form of "providing data, evaluation, and constraints for LLMs."
III. LLMs vs. Deep Learning
1. Deep learning is a much larger family
Deep learning (DL) is a family of techniques using multi-layer neural networks to automatically learn representations from data. Its members include:
| DL branch | Representative | Processing target |
|---|---|---|
| Convolutional networks (CNN) | ResNet, EfficientNet | Images, video, audio |
| Recurrent networks (RNN/LSTM/GRU) | Early machine translation, speech recognition | Sequential data |
| Transformer | GPT, BERT, Llama, Vision Transformer | Primarily text, extended to vision/audio/multimodal |
| Generative models | GAN, diffusion models (Stable Diffusion) | Image and audio generation |
| Reinforcement learning (RL) | AlphaGo, PPO in RLHF | Decision sequences |
LLMs are one application form of deep learning: they use deep learning's tools (neural networks, backpropagation, large-scale distributed training), but they have their own paradigm characteristics — pretrained language modeling + alignment. Conversely, most of deep learning has nothing to do with LLMs: image classification, object detection, speech synthesis, recommendation systems, and RL gaming are not LLMs.
text
Artificial Intelligence (AI)
└─ Machine Learning (ML)
├─ Traditional methods: decision trees, SVM, Bayesian
└─ Deep Learning (DL)
├─ CNN → vision
├─ RNN → sequence
├─ Transformer
│ ├─ Text pretraining → Large Language Models (LLMs) ★ This book's topic
│ ├─ Vision Transformer (ViT) → vision
│ └─ Multimodal Transformer → multimodal large models
├─ Diffusion models → image generation
└─ Reinforcement learning → games, control2. An easily overlooked inheritance point
Almost all of an LLM's "underlying muscle" comes from deep learning's historical accumulation: backpropagation (1980s), word embeddings (2013 Word2Vec), attention mechanism (2014 introduced to machine translation), residual connections and batch normalization (2015 ResNet), sequence modeling (LSTM). LLMs are not an exception to DL; they are the culmination of DL in the direction of "language modeling + ultra-large scale." Understanding this explains why many deep learning-era techniques (learning rate scheduling, gradient clipping, distributed parallelism) are reused intact in large model training. See Pretraining.
IV. LLMs vs. Statistical Language Models
Statistical language models are the direct ancestors of LLMs. Both share the same core goal: estimating the probability of a text sequence. The difference is in implementation and scale:
| Dimension | Statistical language model (n-gram, etc.) | Neural language model / LLM |
|---|---|---|
| Probability source | Count frequency of n-grams in corpus | Neural network's prediction of the next token |
| Generalization | Unseen n-grams get zero probability (needs smoothing) | Words share representations in vector space, enabling generalization |
| Context length | n (usually ≤ 5) | Thousands to hundreds of thousands of tokens |
| Parameters | Count table of vocabulary × n | Billions to hundreds of billions |
| Core capability | Local co-occurrence statistics | Compress world knowledge through the surrogate task of next-word prediction |
The "predicting the next word" idea of statistical language models is fully inherited by LLMs, which is the continuity repeatedly emphasized in Language Modeling and Brief History. The difference is simply: statistical models are "counting," neural models are "understanding" — the former memorizes co-occurrence frequencies, the latter compresses those frequencies into composable knowledge representations.
V. LLMs vs. Foundation Models
"Foundation model" is a concept proposed by the Stanford team in the 2021 paper On the Opportunities and Risks of Foundation Models: models pre-trained on massive data that can serve as starting points for countless downstream tasks. Its relationship with LLMs is genus-species:
text
Foundation Models
├─ Text: Large Language Models (LLMs) — GPT, Llama, Qwen, DeepSeek
├─ Vision: CLIP, SAM, Vision Transformers
├─ Audio: Speech recognition/generation foundation models
├─ Code: Codex, CodeLlama (usually also counted as LLM variants)
└─ Multimodal: GPT-4o, Gemini, LLaVA (see [Multimodal LLMs](/case-studies/multimodal-llm))Key distinction: foundation models are "capability suppliers" (providing general-purpose representations and generation), while LLMs are the most mature branch among them, with language as the core. In everyday speech, the two are often used interchangeably ("foundation model company" usually means an LLM company), but in rigorous discussion, just remember "LLM ⊂ foundation model." The three shared features of foundation models — pretraining, scale, and adaptability — are the source of the "key components" in What Is a Large Language Model.
VI. LLMs vs. Agents: LLM + Tools = Agent System Substrate
This is the most error-prone point in concept clarification: LLMs and agents are not at the same level.
- An LLM is a model: It takes text in and outputs text. It has no "action" capability — it can't browse the web, call APIs, or log actions (unless simulating via text).
- An agent is a system: It uses an LLM as its "brain," with external tools (function calling), memory, and planning loops, forming a "perceive → think → act → observe" closed loop.
text
Typical agent loop (ReAct pattern):
User request ──→ LLM (brain) ──→ Decide to call a tool
↑ │
│ Tool execution (search/code/API)
│ │
└── Observe results ──┘
(Feed tool output back to the LLM, continue reasoning)| Dimension | LLM | LLM-based Agent |
|---|---|---|
| Essence | Single model | Model + tools + memory + control loop |
| Can it act? | Only outputs text | Can call external tools, perform actions |
| Stateful? | No (each call independent) | Yes (has memory: conversation history / long-term memory) |
| Error impact | Wrong output | Wrong actions, potentially cascading consequences |
| Representative | GPT-4, Llama | AutoGPT, Devin, Manus, etc. |
One-line memory trick
"LLMs are the brain; agents are brain + hands + feet + notepad." Without an LLM, an agent has no reasoning core; without tools and loops, an LLM is just a talking model. The full expansion of agents is at Agents with LLMs; its boundary with an "agent handbook," risks (prompt injection, runaway behavior, cost) are specifically addressed on that page.
VII. LLMs vs. RAG Systems
RAG (Retrieval-Augmented Generation) is a system paradigm built around LLMs, not a different type of model: it combines an LLM's generation capability with an external knowledge base's retrieval capability.
text
RAG flow (simplified):
User question ──→ Retriever: Find relevant passages from the knowledge base / vector store
│
└──→ Prompt assembly: question + relevant passages ──→ LLM ──→ Answer with evidence| Dimension | Bare LLM | RAG system |
|---|---|---|
| Knowledge source | Static knowledge learned during pretraining | Dynamically injected external documents / databases |
| Knowledge timeliness | Before the training cutoff date | Index can be updated at any time |
| Traceability | Hard (don't know where knowledge came from) | Can provide retrieval basis (cited sources) |
| Hallucination mitigation | Limited | Significant (answers constrained by retrieved content) |
| Best for | General Q&A, creative writing | Private knowledge, real-time info, citation-required scenarios |
Key clarification: RAG does not change the LLM itself (model weights stay the same). It changes "what input you feed the model." It's a complementary adaptation route to fine-tuning (changing the model). Trade-offs are detailed in RAG and the "long context vs. RAG" trade-off in Context and Long Contexts.
VIII. LLMs vs. AGI: Why "LLM ≠ General Intelligence"
This is the most important and most easily derailing clarification. AGI (Artificial General Intelligence) refers to general intelligence that reaches or exceeds human-level performance on all cognitive tasks. LLMs still have multiple structural distances from it:
| Dimension | LLM status | AGI requirements | Gap |
|---|---|---|---|
| Task scope | Primarily language, extending to multimodal | Language + perception + action + planning + social… | Large (see boundaries in Multimodal LLMs) |
| World model | "Shadow world" implicitly captured in language statistics | Real causal modeling of the physical world | Large |
| Learning method | One-time pretraining + limited continual learning | Lifelong learning, sample-efficient | Large |
| Initiative and goals | No intrinsic goals, only "follows instructions" | Self-set and pursue goals | Large |
| Reliability | Hallucination, instability, uncalibrated confidence | Highly reliable, verifiable | Large |
| Alignment | Helpfulness/honesty/safety still require manual tuning | Internalized value alignment | Unsolved |
Why the illusion that "LLMs are close to AGI"?
Because language is the largest carrier of human thought, a model that excels at language and can connect tools easily creates the illusion that "it has thoughts." But fluent language ≠ understanding (a model can perfectly simulate viewpoints without holding them), and task versatility ≠ generality (it's "pattern matching under sample coverage," not "adaptive intelligence for any novel task"). Treating LLM capabilities as a precursor to AGI is reasonable optimism; treating them as AGI already achieved is a dangerous misjudgment — see Safety and Risks for safety implications.
IX. Why LLMs Became the Synonym for NLP in the 2020s
To close the loop on the opening observation: LLMs "took over" NLP in the 2020s not because they eliminated NLP's problems, but because they provided a unified, scalable solution path — pretraining + alignment + prompt adaptation. This path's three advantages (broad capability coverage, engineering unification, talent and ecosystem concentration) overwhelmed the traditional "task-specific model" route. But this doesn't mean the boundaries disappear:
- At the system level, LLMs are just a component — agents, RAG, evaluation, and deployment are all complete engineering efforts themselves. See Overall Architecture Anatomy.
- At the research level, NLP's open problems (low-resource languages, long documents, faithfulness, interpretability) have not disappeared; they've become new challenges of the LLM era. See Frontier Progress.
- At the capability level, the gap between LLMs and AGI isn't just "make it a bit bigger." It's a qualitative leap in learning methods, world models, and reliability.
A closing diagram
DL ⊃ Transformer ⊃ LLM ⊂ foundation model; LLMs are NLP's dominant implementation; LLM + tools + memory = agent substrate; LLM + retrieval = RAG system; LLMs are one candidate route to AGI. Remember these four sentences, and your conceptual coordinate system is set.
X. Comprehensive Test: Ten Scenarios
Ground the clarification in scenarios. Test whether you've truly distinguished these concepts. First, decide for each scenario which category it belongs to and what technology to use, then check the answer:
| Scenario | Correct classification | Criterion |
|---|---|---|
| Build a sentiment analysis API | NLP task, can use LLM prompting or a dedicated small model | The task definition is NLP; implementation can take many forms |
| Use BERT for text vector retrieval | Encoder model, belongs to deep learning, not LLM (no generation capability) | Look at architecture and training objective, not just "depth" |
| Let the model search the web and answer | RAG or Agent (LLM + retrieval/tools) | External actions = system, not a bare model |
| Let the model plan tasks and operate software autonomously | Agent | Planning loop and tool execution |
| Judge "whether models will eventually surpass humans" | AGI discussion | Involves definitions of general intelligence, not a specific engineering task |
| Use n-gram to analyze text | Statistical language model | Count frequency, no neural network |
| Fine-tune Llama for legal Q&A | LLM + fine-tuning | The model itself is an LLM; fine-tuning is the adaptation method |
| Joint Q&A on images + text | Multimodal foundation model | Beyond pure LLM; see Multimodal LLMs |
| Conversational bot (no external tools) | LLM (conversational form) | No external actions; still a model application |
| Company says it's a "foundation model company" | Foundation model / LLM vendor | Genus-species relationship; colloquial interchangeability is normal |
The takeaway: most scenarios are not "either LLM or something else," but "which layer is LLM at, and what system is it paired with." Being able to accurately say "this scenario = LLM + some system" means you've graduated from concept clarification.
Further Reading
- What Is a Large Language Model — The "ontology" of this page's clarification: three-layer definition and capability inventory
- Brief History — Three paradigm shifts: statistical language model → neural language model → large models
- Language Modeling — The shared "predicting the next word" foundation of LLMs and statistical language models
- Agents with LLMs — The full systematic expansion of "LLMs are the brain"
- RAG: Retrieval-Augmented Generation — The system paradigm of LLM + retrieval, and the trade-offs with fine-tuning / long context
- Safety and Risks — The safety implications of "LLMs are not general intelligence"
References
- Bommasani et al. On the Opportunities and Risks of Foundation Models (2021) — The original source of the "foundation model" concept, defining the three features: pretraining + scale + adaptability
- Brown et al. Language Models are Few-Shot Learners (GPT-3, 2020) — The milestone of "language models as task solvers," the starting point of the few-shot paradigm
- Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models (2022) — Interleaved reasoning + action, the typical paradigm for agent loops
- Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020) — The original RAG paper, the source of "LLM + retrieval"
- Bengio et al. A Neural Probabilistic Language Model (2003) — The pioneering work of neural language models, the watershed between LLMs and statistical language models
- OpenAI · On the Definition and Discussion of AGI (official documentation) — Using OpenAI's charter's cautious definition of AGI as an example, illustrating how the industry defines "general intelligence"