Appearance
AI vs ML vs DL vs GenAI vs Agents
In one sentence: AI is the goal, machine learning is the means, deep learning is a branch of machine learning, generative AI is the subset of deep learning that aims to produce new content, and agents are an application paradigm built on top of LLMs. These five terms are not parallel to one another; they form a layered structure—nested within each other, yet each with its own identity.
You can scroll past "AI will replace programmers," "machine learning jobs are shrinking," and "the year of the LLM agent" in a single day—three stories about five different things in the same world. Drawing these boundaries clearly is the first gate to understanding today's AI hot topics, and it is the "boundaries" companion to this handbook's big-picture overview.
1. One Overview Table: Who Contains Whom
The conclusion first; the rest of the page expands on it:
| Concept | Full Name | One-Line Definition | Containment | Examples |
|---|---|---|---|---|
| AI | Artificial Intelligence | The overall goal of making machines exhibit intelligence | The largest set; contains all the others | Chess programs, expert systems, self-driving |
| ML | Machine Learning | The path of learning patterns automatically from data | ML ⊂ AI | Recommenders, spam filters, XGBoost |
| DL | Deep Learning | Learning representations with multi-layer neural networks | DL ⊂ ML | CNNs, RNNs, Transformers, LLMs |
| GenAI | Generative AI | Models whose goal is to generate new content | GenAI ⊂ DL | ChatGPT, Midjourney, Sora |
| Agent | Agent | An autonomous execution system with an LLM as its brain | An application paradigm spanning all the layers above | Manus, Deep Research, AutoGPT |
Three containment facts to memorize
AI ⊃ ML ⊃ DL; GenAI is a subset of DL; and an agent is not a new model but an application paradigm built on LLMs. Every clarification below starts from these three sentences.
2. The Five Concepts, One by One
1. Artificial Intelligence (AI): The Big Umbrella Born in 1956
| Dimension | Details |
|---|---|
| Origin | The 1956 Dartmouth Conference, where John McCarthy coined "Artificial Intelligence" and AI was born as a discipline |
| Core problem | Getting machines to do tasks that require human intelligence: reasoning, perception, language, planning, learning |
| Classic schools | Symbolism (rules/logic), connectionism (neural networks), behaviorism (perception–action) |
| Non-ML branches | Search and game playing, planning, knowledge graphs, expert systems, constraint solving |
| Today's usage | In the media, "AI" now almost always means ML/DL-driven systems, but academically the umbrella is much bigger |
The details: AI is the only one of the five terms defined by a goal rather than a method. For two decades after Dartmouth, the mainstream line was symbolic AI: encode knowledge as explicit rules and let machines reason with logic. The medical diagnosis system MYCIN and IBM's Deep Blue beating Kasparov relied on search and rules, not learning. Knowledge graphs are a major product of that tradition—representing entity relationships explicitly with nodes and edges—and they remain a key complement for improving LLM factuality today (see Knowledge Graphs and Knowledge Injection).
Key point: not all AI is machine learning. A chess program using minimax search is AI but not ML; a rules-based approval workflow is AI but not ML. ML is simply the currently most successful of the many paths to machine intelligence. For the full AI lineage, read What Are AI Hot Concepts? and A Brief History.
2. Machine Learning (ML): Learning from Data
| Dimension | Details |
|---|---|
| Definition | Computers discover patterns from data automatically and use them for prediction and decisions; the rules come from data, not from human hands |
| Three paradigms | Supervised (labeled data), unsupervised (unlabeled), reinforcement (reward signals) |
| Relation to AI | ML ⊂ AI; the largest subfield today |
| Relation to DL | DL ⊂ ML; classical methods (tree models, SVMs, linear models) are still in heavy service |
| Strongholds | Tabular data, recommendation ranking, risk control, advertising, search ranking |
The details: the whole secret of machine learning lies in "where does knowledge come from"—not from programmers writing rules, but from models fitting rules out of data. The three paradigms differ by the form of the answer: supervised learning has ground-truth labels, unsupervised learning has only the structure of the data itself, and reinforcement learning has only reward signals after the fact. Far from obsolete, this framework remains the most dependable production toolkit outside large models: e-commerce recommendations, feed ranking, fraud and risk control, and ad click-through prediction are still battlefields where tree models and deep models fight side by side (see Recommender Systems in the LLM Era).
A frequently underestimated fact: on tabular data, classical methods like XGBoost/LightGBM still often beat deep learning. ML is not an "outdated" technology, and DL is not "better ML"—they are two kinds of tools for different shapes of data. Deep learning is a subset of ML; details in the next section.
3. Deep Learning (DL): The Multi-Layer Neural Network Family
| Dimension | Details |
|---|---|
| Definition | Representation learning with multi-layer neural networks: shallow layers learn edges and strokes, deep layers learn object semantics |
| Breakout moment | 2012: AlexNet slashed ImageNet top-5 error from 26.2% to 15.3% |
| Main families | CNNs (vision), RNNs (sequences), Transformers (attention, 2017) |
| Relation to ML | DL ⊂ ML—"ML on a different kind of data" |
| Costs | Data-hungry, compute-hungry, hard to interpret |
The details: deep learning is a "revolution" because it turned ML's feature engineering into representation learning—instead of hand-designing features, the network learns usable representations from raw pixels or characters, layer by layer. AlexNet lit the first fuse in 2012; the 2017 paper Attention Is All You Need introduced the Transformer and lit the second, ultimately giving birth to the LLM era. For the mechanics and trade-offs of CNNs/RNNs/Transformers, see Transformer and Attention; for the full arc of this technical line, see A Brief History.
"Uses neural networks" is not the same as "deep learning"
The key to deep learning is depth—the layers. A single-layer perceptron (1960s) is a neural network but not a deep model; the fact that Minsky proved in 1969 it could not represent XOR is a direct cause of the first AI winter. Depth is the essence of DL.
4. Generative AI (GenAI): Built to Generate
| Dimension | Details |
|---|---|
| Definition | Models whose goal is to generate new content (text / images / audio / video / code) |
| Core difference | Discriminative models learn "what is this"; generative models learn "how to make it" |
| Text workhorse | Large language models (LLMs): Transformer-based pretraining plus alignment |
| Vision workhorse | Diffusion models: progressively reconstructing images from noise |
| Relation to DL | GenAI ⊂ DL—a classification by goal, not by architecture |
The details: generative AI is not a new architecture but a new class of goals. Traditional deep learning models are mostly discriminative—image in, "cat or dog" out. A generative model instead learns the data distribution and then samples new examples from it. Today's two workhorses: LLMs generate text autoregressively with Transformers (see Large Language Models), and diffusion models turn noise into images and video step by step (see Diffusion Models and Generative AI).
The commercialization speed of this subset is unprecedented: ChatGPT reached 100 million users in two months (see ChatGPT and Conversational AI), Midjourney reshaped the illustration industry (see Midjourney and Image Generation), and Sora opened the door to video generation (see Sora and Video Generation). GenAI is the hottest subset of DL right now, but DL is far more than generation.
5. Agents: An Application Paradigm on Top of LLMs
| Dimension | Details |
|---|---|
| Definition | An execution system with an LLM as its "brain," autonomously closing the perceive–plan–act loop |
| Key components | Planning + tool calling + memory + action |
| Relation to models | Not a new model family; the underlying engine is still an LLM—an application paradigm for LLMs |
| Run loop | Receive a goal → break it into tasks → call tools → observe results → iterate until done |
| Examples | Manus, Deep Research, various coding agents |
The details: an agent is the only one of the five terms that names a system rather than a model. It upgrades the LLM from a question-answering machine to an execution machine: give it a goal ("pull the earnings report and write a summary"), and the agent plans tasks, searches the web, invokes code tools, verifies repeatedly, and delivers a result. Its loop is essentially the assembly of five pieces—LLM + planning + tools + memory + action. For the mechanics see AI Agents, for shipping examples see Manus and Agent Apps, and to build one see Build an Agent from Scratch.
Key point: an agent is not a fifth kind of "model"—it's a fourth way of "using" one. What runs underneath is still an LLM; swap in a stronger model and the agent gets stronger; there is no "brain" independent of the LLM. Mistaking agents for a new model class is the source of 90% of the concept confusion out there.
3. The Stack View: From Application to Infrastructure
Shift the perspective from "horizontal containment" to "vertical layering," and the five concepts land on four layers of the technology stack:
┌──────────────────────────────────────────────────────────┐
│ ① Application layer Agents · RAG · Prompt apps │
│ (assembling model capability into products: Q&A, search, automation) │
├──────────────────────────────────────────────────────────┤
│ ② Model layer LLMs · Multimodal · Diffusion models │
│ (GenAI's "generation" happens at this layer) │
├──────────────────────────────────────────────────────────┤
│ ③ Architecture layer Transformers (attention) │
│ (CNNs/RNNs were the old workhorses; Transformers are the modern foundation) │
├──────────────────────────────────────────────────────────┤
│ ④ Infrastructure layer GPUs · Vector databases · Inference engines │
│ (compute, retrieval, acceleration—how fast the three layers above can run) │
└──────────────────────────────────────────────────────────┘- Application layer: agents, RAG, and prompt apps sit closest to the user. RAG fixes the LLM's hallucination and knowledge-staleness problems (see Retrieval-Augmented Generation (RAG) and Build a RAG App from Scratch); prompt engineering sets the floor for output quality (see Prompt Engineering and Prompt Playbook); agents assemble these capabilities into autonomous products (see AI Agents).
- Model layer: LLMs carry text ability (see Large Language Models), multimodal models read text, images, audio, and video together (see Multimodal Models), and diffusion models handle image and video generation (see Diffusion Models and Generative AI).
- Architecture layer: the foundation of nearly every frontier model today is the Transformer (see Transformer and Attention). This layer sets the ceiling for models and is the main battleground of paper reading.
- Infrastructure layer: vector databases power semantic retrieval (see Vector Databases and Semantic Search), and inference engines plus quantization determine deployment cost (see Inference Optimization and Quantization). Beginners often overlook this layer, yet it decides whether a product can ship. For the full stack, see Overall Architecture Dissected.
Use these four layers to frame any AI news
Next time you see an AI headline, ask first: which layer is the breakthrough in? A model-layer breakthrough (a new model) propagates up to the application layer, but an application-layer innovation (an agent) does not need a new model. Layered thinking filters out 80% of useless anxiety.
4. Four Common Confusions
| Confusion | The Truth | One-Line Verdict |
|---|---|---|
| AI = LLM? | Large models are just AI's current star subset | AI is a century-old umbrella discipline; the LLM is the newest, sharpest blade under it |
| ML = DL? | DL is just the deep subset of ML | Tree models still often win on tabular data; DL is not "stronger ML" but "another kind of ML" |
| Generative = intelligent? | Generation ability ≠ reasoning, understanding, reliability | Writing a decent essay is not the same as doing arithmetic, not hallucinating, or being accountable |
| Agent = new model? | An agent is a way of using an LLM, not a new "brain" | The same LLM runs underneath; swap the model and you swap the agent's brain |
One by one:
"AI = LLM?" In media usage the two are nearly interchangeable, but academically and in engineering they differ by an order of magnitude. AI includes symbolic AI, knowledge graphs, search and planning, robotics, and dozens of other branches; the LLM is only the latest link in the chain DL → Transformer → pretrained large models. Equating AI with LLM is like equating "transportation" with "cars"—you can communicate that way, but it isn't precise.
"ML = DL?" Deep learning is a branch of machine learning, but machine learning is much more than deep learning. In large-scale production systems—recommenders, credit risk, ad auctions—tree models and classical methods still hold the main-force positions (see Recommender Systems in the LLM Era). The correct statement: DL dominates on unstructured data; classical ML still holds its own on tabular data.
"Generative = intelligent?" This is the most dangerous of the four confusions. Generation ability (writing fluent prose) and intelligence (reasoning correctly, honestly admitting ignorance, reliably taking responsibility) are different things. LLMs will fabricate facts with a straight face (hallucination)—which is exactly why RAG and alignment exist, and exactly what LLM evaluation and benchmarks are for. Being able to write is not being able to think, let alone being trustworthy.
"Agent = new model?" An agent's intelligence comes entirely from the underlying LLM; it has no model parameters of its own. The difference is only in "how you use it": treat it as a Q&A box and it's a chatbot; give it planning, tools, and memory and it's an agent. For a more systematic treatment, see AI Agents.
A word of warning
These confusions are not word games. In an interview, describing "an agent is an application paradigm" as "an agent is a new model," or saying out loud that "the LLM is AI," can sink you in one stroke—concept boundaries are table stakes in this field.
5. The Job-Hunting View: What Each Term Means in a Job Description
The same JD points to five different skill stacks depending on the keyword. Cross-reference the JD Knowledge Map and the JD List:
| JD Keyword | Actual Skill Stack | Typical Roles |
|---|---|---|
| AI / artificial intelligence | A generalist term, often spanning the full algorithm stack: math + at least one technical line + paper reading | AI scientist, AI researcher |
| Machine learning | Feature engineering, tabular modeling, tree models/GBDT, model evaluation, SQL/Spark | ML engineer, algorithm engineer (risk/recommendations/ads) |
| Deep learning | PyTorch, CNNs/Transformers, training and tuning, GPU resource management | Algorithm engineer (CV/NLP), deep learning researcher |
| Generative AI / LLM | Prompt engineering, RAG, fine-tuning (LoRA), evaluation, agent development | LLM application engineer, prompt engineer, LLM algorithm engineer |
| Agent | Engineering skills: tool calling, MCP, workflow orchestration, frontend/backend integration | Agent application developer, AI product engineer |
A counterintuitive observation: the closer a role is to the application layer, the higher the engineering share and the lower the modeling depth. The fastest-growing role after 2024 isn't "LLM researcher" but "LLM application engineer"—JDs that mention "agent" or "RAG" are testing engineering and system-design ability. To benchmark your own skills, see Resume Analysis: What to Highlight; to test your grasp of concepts, see the Interview Question Bank.
The minimal action plan
Don't be intimidated by five terms. Master one main line and let the "containment + four layers" picture carry the rest: memorize this page's tables, then follow Learning Paths starting from how large models work, and that covers 90% of interview concept questions. For quick lookups, return to the Glossary anytime.
Further Reading
- What Are AI Hot Concepts? — the companion "big picture" page; the two are two sides of one map
- A Brief History — 1956 to 2025, how five lines emerged and succeeded one another
- Overall Architecture Dissected — the complete stack from GPUs to agents
- Large Language Models — the internals of GenAI's text workhorse
- AI Agents — the LLM + planning + tools + memory + action loop, dissected
- Build an Agent from Scratch — turn the "application paradigm" into code
- Retrieval-Augmented Generation (RAG) — the engineering pattern that fixes hallucination and stale knowledge
- The Paper Map — when you want to dig deeper, revisit everything from the papers' angle
- Glossary — more confusable term pairs beyond these five
References
- Dartmouth Workshop proposal (1955, McCarthy et al.) — the founding document of the term "Artificial Intelligence"
- Russell & Norvig, Artificial Intelligence: A Modern Approach — the discipline's most authoritative textbook and a complete map of AI's scope
- McCulloch & Pitts, A Logical Calculus of the Ideas Immanent in Nervous Activity (1943) — the earliest mathematical model of neural networks
- Rosenblatt, The Perceptron (1958) — the single-layer perceptron and the first boom
- Krizhevsky et al., ImageNet Classification with Deep Convolutional Neural Networks (NeurIPS 2012) — AlexNet, the spark of the deep learning revival
- Vaswani et al., Attention Is All You Need (NeurIPS 2017) — the Transformer, foundation of the modern architecture layer
- Ho et al., Denoising Diffusion Probabilistic Models (NeurIPS 2020) — diffusion models, the source of the ideas behind image/video generation
- Brown et al., Language Models are Few-Shot Learners (NeurIPS 2020) — GPT-3, the emblem of LLM scaling laws
- Doshi-Velez & Kim, Towards A Rigorous Science of Interpretable Machine Learning (2017) — model interpretability and the "black box" debate