Appearance
What Is an AI Agent
This page is the thesis statement of the entire site. Every chapter that follows — Agent Loop, tools, memory, planning, multi-agent systems, evaluation — is an expansion of the definition laid out here. So it's worth reading slowly: nail down "what an agent is" first, then talk about how to learn the rest.
One detour before we start: "agent" is an old word in AI. Reinforcement learning and multi-agent systems research had been using it for decades. After 2023 it was reinvented — the core shifted from "an entity that learns a policy in an environment" to "a system that uses an LLM to make decisions." This site is about the latter: what the industry now means by default when it says "Agent." The two are conceptually related, but their technology stacks are almost entirely different, so watch the context when reading older literature.
1. The One-Sentence Definition
Here is the working definition this site adopts:
An AI Agent (LLM Agent) is a system driven by a large language model that autonomously uses tools to complete multi-step tasks in an environment.
Five key phrases in that sentence, none of them optional:
- LLM-driven: the decision core is a language model, not a rules engine. This is the dividing line between post-2023 Agents and the classical AI notion of an "agent."
- Autonomous: the model decides what to do next itself, rather than following a flow hard-coded by a human.
- Uses tools: it can call functions, query databases, send requests, read and write files — it can genuinely affect its environment.
- In an environment: tool results feed back in, and the model corrects its behavior based on "what the environment actually said," rather than generating in a vacuum.
- Multi-step tasks: a single question-and-answer exchange is not an Agent. An Agent's value shows up precisely in tasks that take ten or twenty steps to finish.
How the Major Definitions Compare
There is no single canonical answer to "what is an Agent" in industry. Anthropic concedes this in the opening of Building Effective Agents: some people use the word for fully autonomous systems, others for implementations that follow predefined flows. They lump these together as agentic systems and draw one crucial architectural line:
"Workflows are systems where LLMs and tools are orchestrated through predefined code paths. Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks." — Anthropic, Building Effective Agents (December 2024, Erik Schluntz & Barry Zhang)
In plain terms: a workflow is a system where LLMs and tools are orchestrated along predefined code paths; an agent is a system where the LLM dynamically directs its own process and tool usage. The essence of the difference is who owns the control flow — a human-written if/else, or the model's own judgment.
OpenAI, in A Practical Guide to Building Agents (April 2025), offers a more product-flavored definition:
"Agents are systems that independently accomplish tasks on your behalf."
And then adds a sharp caveat: systems wired to an LLM but not letting the LLM control execution — simple chatbots, single-turn Q&A, sentiment classifiers — are not Agents. OpenAI also splits an Agent into three core components: Model (the LLM that drives reasoning), Tools (callable external functions/APIs), and Instructions (behavioral rules and guardrails).
IBM's definition is more restrained: "An AI agent is a software program capable of acting autonomously to understand, plan and execute tasks" — software that can autonomously understand, plan, and execute tasks, powered by an LLM and able to connect to external tools and data sources.
| Source | Core claim | Emphasis |
|---|---|---|
| Anthropic | LLM dynamically directs its own process and tool usage | Architectural distinction: workflow vs. agent |
| OpenAI | A system that independently accomplishes tasks on your behalf | Product view: taking over the whole workflow |
| IBM | Software that autonomously understands, plans, and executes | The classical agent concept, LLM-ified |
The wording differs, but all three point to the same structural test: who is making the decisions, and are those decisions updated in the loop based on environmental feedback? That is also this site's criterion for "is this system an Agent"; see Concept Clarification for details.
A frequent interview question
"What's the difference between an agent and a workflow?" is nearly a guaranteed question for Agent roles. The standard answer is Anthropic's line: systems whose control flow is governed by predefined code are workflows; systems governed dynamically by the model are agents. But be ready to keep going — the two are a spectrum rather than a binary (Section 3 below), and the overwhelming majority of production systems are hybrids of both.
2. The Agent Formula, Unpacked
The most widely circulated engineering formula in the field is:
Agent = LLM (the brain) + Planning + Memory + Tool UseThis decomposition first appeared systematically in Lilian Weng's "LLM Powered Autonomous Agents" and the Fudan NLP team's survey "A Survey on Large Language Model based Autonomous Agents," and has since become the de facto analytical framework. As a picture:
┌─────────────────────────────┐
│ User Task │
└──────────────┬──────────────┘
▼
┌─────────────────────────────────────┐
│ LLM (Brain / Reasoning Core) │
│ Understand the goal · Decide the │
│ next step · Emit an action │
└──────┬───────────────┬──────────────┘
│ │
┌──────────▼────┐ ┌──────▼───────────┐
│ Planning │ │ Memory │
│ Task breakdown│ │ Short: context │
│ Reflection & │ │ Long: vector DB │
│ self-correction│ │ and more │
└───────────────┘ └──────────────────┘
│ │
▼ ▼
┌─────────────────────────────┐
│ Tools (Tool / Execution │
│ Layer) │
│ Search · Code exec · Files │
│ · API │
└──────────────┬──────────────┘
▼
┌─────────────────┐
│ Environment │ The environment returns
│ (env. feedback) │──┐ real results
└─────────────────┘ │ (Observation)
▲ │
└── Feedback to the LLM, next loop iterationWhat problem does each of the four components solve?
- The LLM is the brain, but it isn't everything. A bare model can only "talk," not "do." The essence of agent engineering is converting a model's language ability into decision-making ability.
- Planning solves "one step isn't enough to think it through": task decomposition, stepwise reasoning, and post-execution reflection and self-correction. See Planning & Reasoning.
- Memory solves "the context can't hold it all, and nothing survives across sessions": short-term memory is just the conversation history inside the context window; long-term memory usually lives in external storage plus retrieval. See Memory Systems.
- Tools are the agent's hands and eyes: without them, the model's only influence on the world is emitting text. The quality of tool calling (descriptions, parameter design, error messages) often decides an agent's ceiling more than the model choice does. See Tools & MCP.
Anthropic calls the combination of these three capability boosters the augmented LLM — a model call equipped with retrieval, tools, and memory, the basic building block of every agentic system. The framing is worth remembering: building an agent isn't building a rocket from scratch; it's wiring four mature components together correctly. The engineering difficulty has never been in any single component but in their seams: how to write tool descriptions so the model doesn't pick the wrong one, whether retrieved memories will pollute the context, how to detect and recover when planning fails. The five articles in this site's Core Components section are, at heart, a tour of those seams.
Anthropic's minimalist view
Alongside that definition, Anthropic puts it even more bluntly: "agents are typically just LLMs using tools based on environmental feedback in a loop" — an Agent is usually just "an LLM using tools in a loop, based on environmental feedback." Planning and memory are not mysterious standalone modules: planning is the reasoning in the prompt, and short-term memory is a message list that keeps growing. The right way to start is to get this minimal loop running and add what's missing, rather than reaching for a heavyweight framework on day one.
3. The Autonomy Spectrum: An Agent Is Not a Binary Concept
The argument beginners most easily fall into is "does this system count as a real Agent?" The question itself is malformed. Anthropic files both workflows and agents under agentic systems, which hints at a more important fact: autonomy is a continuous spectrum, not a black-and-white boundary.
Low ◄──────────────────────── Autonomy ────────────────────────► High
Single LLM call Fixed workflow Constrained agent Fully autonomous agent
(QA/translation) (chaining/routing/ (humans gate the (long stretches of
parallelization) critical checkpoints) unsupervised execution)
│ │ │ │
│ │ │ │
Model only Model follows a Model drives the Model plans, executes,
emits text human-written flow, but steps/ and self-corrects,
path and never permissions are runs unattended
chooses the next limited, humans
step can take overWhere a system sits on this spectrum is determined by three variables:
- Who owns the control flow: the more of it lives in predefined code paths, the further left; the more the model decides, the further right.
- How big the loop is: a single call sits at the far left; a tool loop with a step cap and exit conditions sits in the middle; something that could in principle run forever sits at the far right.
- Reversibility of actions: read-only tools (search, query) are safe to push rightward; write operations (sending email, modifying code, payments) usually need to be pulled back to the middle with human-in-the-loop gates — see Human-in-the-Loop.
The practical conclusion is clear: the overwhelming majority of production systems should be built in the middle of the spectrum, not at the right end. Anthropic's advice is "find the simplest solution that works, and add complexity only when necessary" — if a single call plus retrieval gets it done, don't build a workflow; if a workflow gets it done, don't build an agent. OpenAI's guide likewise recommends maximizing a single agent's capability before considering multi-agent. This isn't conservatism; it's scar tissue. Every step to the right raises cost, latency, and error compounding — see Cost & Latency.
Three Judgment Calls to Test Your Understanding
- "A customer-service bot wired to an LLM that replies along a fixed script tree — is that an Agent?" No. The control flow is the script tree (predefined code paths); the LLM only polishes wording. It's a workflow — arguably not even that. OpenAI explicitly excludes this category.
- "A script that calls an LLM in sequence to summarize, translate, and proofread — is that an Agent?" No, that's a prompt-chaining workflow. What each step does and in what order is entirely hard-coded by a human. It may be very useful, but the model has zero autonomy.
- "A loop where the model decides what to search, what to read, and when to stop, but only read-only tools are allowed — is that an Agent?" Yes — an agent with constrained autonomy. Restricting tool permissions doesn't change who owns the control flow; this is exactly the typical middle-of-the-spectrum shape.
If you got all three right, you've passed the definition; if not, go back to Section 1 and reread Anthropic's two sentences.
4. Why Now: The Technical Preconditions of the Agent Boom
The agent concept has existed in AI for decades, but it only became genuinely usable in 2023. That's not a marketing cycle — several critical puzzle pieces fell into place between 2022 and 2025:
| Date | Event | Which piece it supplied |
|---|---|---|
| Oct 2022 | ReAct paper published (arXiv:2210.03629) | Alternating "reasoning-action" traces via prompt proved LLMs can think and act simultaneously |
| Jun 2023 | OpenAI ships function calling | Tool calling went from the engineering acrobatics of "parsing the model's text output" to a structured API — a step change in reliability |
| Feb 2024 | Gemini 1.5 Pro announced, million-token context in preview | Long context lets a multi-step task's history, documents, and tool results fit in one window |
| Sep 2024 | OpenAI releases o1-preview | Reasoning models arrive: the model learns to "think a while longer" before answering; complex planning jumps |
| Oct 2024 | Anthropic ships computer use (beta) for Claude 3.5 Sonnet | For the first time a model can look at the screen and drive the mouse and keyboard; the tool surface extends from APIs to any GUI |
| Nov 2024 | Anthropic open-sources the Model Context Protocol (MCP) | Tool/data-source integration gets one open standard; the N×M integration problem becomes N+M |
| Jan 2025 onward | DeepSeek-R1 and other open reasoning models crash the price of reasoning | Token costs for long-loop agents enter affordable territory |
| Mar 2025 | Manus launches and goes viral; general-agent products break into the mainstream | "The year of the agent" is confirmed from the product side; capital and hiring markets follow |
Viewed another way, this table answers "when did each of the agent's four components become usable":
- Tool use: function calling solved "call reliably" in 2023; MCP solved "integrate in a standardized way" at the end of 2024.
- Memory/context: million-token context windows plus mature RAG engineering let agents handle real-world-scale material.
- Planning/reasoning: the reasoning-model paradigm o1 started (RL-trained test-time compute) took multi-step planning from "frequently falls over" to "broadly usable."
- Environment interaction: computer use freed agents from the "must have a ready-made API" constraint — they can now operate any software interface.
As of 2026, these four pieces are no longer research problems; they're engineering problems. That is the essential difference between "learning agents now" and "learning agents three years ago." For a fuller chronology, see A Brief History of Agents.
Don't get swept up by "the year of the agent" narratives
Somebody declares the year of the agent annually. The technical puzzle pieces really did come together in 2023–2025, but "usable" is not "reliable." Today's agents already create real value in structured, verifiable domains (coding, retrieval, data manipulation); in open-ended, irreversible, high-stakes scenarios they still need plenty of human backstopping. Think of them as a highly capable intern who makes rookie mistakes — that's closer to reality than "digital employee."
5. A Minimal Agent: 30 Lines of Pseudocode
Once the definition is clear, the single most useful thing to internalize is that the agent's core loop is almost disappointingly simple. This pseudocode (Python-flavored, with API calls simplified) is a fully functional agent:
python
# Minimal agent: the model decides in a loop whether to call a tool or give the final answer
def run_agent(task, tools, max_steps=20):
messages = [
{"role": "system", "content": "You can use tools to complete the task. Think before you act."},
{"role": "user", "content": task},
]
for step in range(max_steps):
# 1. The model sees the full history and decides: call a tool or answer directly
response = llm.chat(messages, tools=tools.schemas)
# 2. The model chose not to call tools → task done, return the final answer
if not response.tool_calls:
return response.content
# 3. Execute every tool call the model requested and append the results to the context
messages.append(response)
for call in response.tool_calls:
try:
result = tools.execute(call.name, call.arguments)
except Exception as e:
result = f"Tool execution failed: {e}" # feed errors back too; let the model self-correct
messages.append({"role": "tool", "tool_call_id": call.id,
"content": str(result)})
return "Max steps reached; task incomplete" # always have an exit conditionUnpacking this code, it contains every essential element of an agent:
- A loop plus an exit condition: the
forloop is the Agent Loop;max_stepsis the guardrail that keeps a runaway loop from burning money. Real systems add more exit conditions: timeouts, token budgets, human takeover, and so on. - The model owns the control flow: what happens next is decided entirely by the output of
llm.chat— exactly the code-level form of Anthropic's "LLM dynamically directs its own process." - A closed feedback loop with the environment: tool results (including errors) are appended to
messages, so on the next turn the model sees the real consequences and can self-correct. This is the engineering realization of the ReAct paper's alternating "reasoning + acting" cycle. - Memory is just
messages: short-term memory has no separate module — it's that ever-growing list. Only when it outgrows the window do you need summarization, compaction, and external retrieval — which is what context engineering is for.
Every framework on the market — LangGraph, the OpenAI Agents SDK, the Claude Agent SDK — is essentially state management, streaming output, interrupt/resume, and permission control wrapped around this loop. Hand-write the loop first, then pick a framework, and you'll know exactly what each layer of abstraction is doing for you. For a hands-on walkthrough, see Build Your Own Agent from Scratch and the step-by-step tutorial.
6. Capability Boundaries: What Agents Can and Cannot Do
Setting the right expectations matters more than learning to write agents. Here is an honest assessment as of 2026.
Already Reliable
- Coding tasks: code has automated tests as objective verification, and environmental feedback (compile errors, test results) is crisp — the first domain agents cracked. Claude Code and Cursor are both built on this bet; see the Claude Code case and Cursor case.
- Retrieval and research tasks: clear goals, a "search–read–synthesize" loop, high error tolerance (if a search misses, search again). Perplexity and Deep Research–style products run at scale; see the Perplexity case.
- Automation of structured flows: support tickets, data entry, report generation. OpenAI's guide names three high-value categories: complex decisions, rules too hard to maintain, and heavy dependence on unstructured data.
Still Hard
- Error compounding over long chains: 95% per-step accuracy is about 36% after twenty steps. Multi-step tasks are where agents shine — and their Achilles' heel: the more steps, the more you need evaluation and guardrails. See Evaluation.
- Irreversible, high-risk actions: payments, deletions, outbound messages. A single hallucination can cause real damage; these actions need human confirmation or hard-coded rules as a backstop. See Security & Guardrails.
- Judging fuzzy goals: tasks like "just make it better" with no acceptance criteria — the agent will burn its budget "looking busy." Successful agent tasks almost always have a crisp success criterion.
- Long-term memory across sessions: current approaches (RAG, summaries, structured profiles) are patches, not solutions. Agents will "forget" what you taught them last week.
- Cost and latency: a twenty-minute multi-step task can burn millions of tokens. If one call can do it, don't make twenty — worth taping to your monitor.
One-line summary of the boundary: agents excel at tasks that are "clearly specified, verifiable, and safe to retry"; they struggle with tasks that are "fuzzy, irreversible, and unforgiving." Ask these three questions before picking a scenario — it beats agonizing over framework choice. For more cautionary tales, see Common Pitfalls.
7. Where to Go Next
The whole site is organized around this page's definition. Suggested order:
- See the whole picture: read Agent System Anatomy for the complete component diagram of a real agent system; then Concept Clarification to separate agent, workflow, RAG, Copilot, and friends.
- Master the core loop: start with the Agent Loop, then read Tools & MCP, Memory, Planning, and Context Engineering. Together these five components are the full expansion of this page's formula.
- Build one: follow Build Your Own to hand-write a framework-free agent, then rebuild it with a framework (see Framework Guide) and feel what each abstraction layer costs and buys.
- Advanced architecture: only after a single agent feels comfortable, touch multi-agent systems, evaluation, and observability — keep the order, because multi-agent is over-engineering most of the time.
- Study real products: read the case-study section (starting with Claude Code) for teardowns of Claude Code, Manus, Devin, and others, and place each on this page's autonomy spectrum.
- For the job hunt: if you're aiming to switch into the field, go straight to The Job Landscape and the Knowledge Map and work backward from job descriptions.
For the full route design, see Learning Paths.
References
- Anthropic: Building Effective Agents — the canonical workflow-vs-agents distinction; the first must-read of agent engineering (December 2024).
- OpenAI: A Practical Guide to Building Agents — OpenAI's agent-building guide; introduces the Model/Tools/Instructions triad (April 2025).
- Anthropic: Introducing the Model Context Protocol — the official MCP announcement (November 2024); the starting point for understanding tool-integration standardization.
- ReAct: Synergizing Reasoning and Acting in Language Models — arXiv:2210.03629, the foundational paper of the alternating "reasoning + acting" loop (October 2022).
- IBM: AI Agents in 2025 — Expectations vs. Reality — IBM's definition of AI agents and analysis of capability boundaries.
- Google: Introducing Gemini 1.5 — the opening of the million-token context era (February 2024).