Appearance
Concepts Clarified: Agent / Workflow / Copilot / RAG
This industry's terminology was wrecked jointly by marketing departments and researchers. The "Agent" one company shipped last week might be nothing but three prompts wired in sequence; meanwhile another team's "workflow" actually lets the model decide every step on its own. Muddled concepts aren't just embarrassing — they directly distort technology selection, architecture, and interview answers. An engineer who dresses up a RAG retrieval-QA system as a "multi-agent platform" and an engineer who can recite definitions but can't tell routing from an agent have fallen into the same hole.
This article has one aim: give each high-frequency concept an operational definition, plus typical examples and a "tell," so that you can classify any system within ten seconds. A selection decision flowchart closes the piece.
1. The Baseline Definition: What Exactly Is an Agent
Anchor first. The two most influential definitions in the industry:
- Anthropic (Building Effective Agents, December 2024 — now the widely cited engineering guide) groups all related systems under agentic systems and draws one architectural line inside them:
- Workflow: systems where LLMs and tools are orchestrated through predefined code paths.
- Agent: systems where the LLM dynamically directs its own process and tool usage, controlling for itself how the task gets done.
- OpenAI (A Practical Guide to Building Agents) defines it more from the product angle: an Agent is a system that independently accomplishes tasks on your behalf.
Put the two together and three necessary conditions for an agent emerge:
- An LLM is in the decision loop — not a rules engine, not a fixed script;
- The model decides the next step at runtime — which tool to call, with what arguments, when to stop; not pre-written by the developer;
- Goal-oriented, not single-turn response — it works toward a "task" and may run across many turns, many tools, and long stretches of time.
To judge whether a system is an agent, the core test is condition 2: who makes the control-flow branching decisions. Code decides → workflow. Model decides → agent. This test runs through the whole article.
┌─────────────────────────────────────────┐
│ Agentic System │
│ (the umbrella term for every system │
│ that "gets an LLM to do work") │
├───────────────────┬─────────────────────┤
│ Workflow │ Agent │
│ Control flow │ Control flow │
│ lives in code │ belongs to model │
│ Predictable, │ Flexible, │
│ reproducible │ autonomous, costly │
└───────────────────┴─────────────────────┘A quick mental model
A workflow is a musical score and the LLM is the musician playing what's written; an agent is jazz — you hand over a theme and rules (the goal, the tools, the constraints), and where it goes is decided live by the musician. Their cost structures, failure modes, and testing approaches are completely different — which is exactly why you must distinguish them first.
2. Agent vs. Workflow: the Most Valuable Line
This is the most valuable pair to distinguish, because it maps directly onto the engineering decision "does this requirement need an agent framework?"
Workflow: the flowchart is drawn by a human; the LLM is merely the executor of certain nodes.
- Typical example: a support-ticket pipeline — step one, classify (LLM); step two, route by category to a template reply (if-else code); step three, sensitive-word screening (code); step four, send. The model never once decides "what happens next."
- Anthropic's article catalogs five common workflow patterns: prompt chaining, routing, parallelization, orchestrator-workers, and evaluator-optimizer. Note: even patterns like "routing" that appear to "make decisions" have a predefined decision space — the system can only pick from routes A/B/C, never invent a fourth.
Agent: the developer provides the goal, the toolset, and the guardrails; how many loop iterations run and what each one does is the model's call. Details in the Agent Loop.
- Typical example: tell a coding agent "fix the failing tests in this repo" — it decides which file to read first, what commands to run, what to change, then re-verifies, looping until tests pass or it gives up. Claude Code and SWE-Agent are both this kind.
Side by side:
| Dimension | Workflow | Agent |
|---|---|---|
| Control flow | Predefined by the developer | Decided by the model at runtime |
| Predictability | High; behavior enumerable | Low; trajectories never repeat |
| Latency/cost | Low and stable | High and volatile (loop count unknown) |
| Suited tasks | Steps known, paths enumerable | Step count and path unknowable in advance |
| Testing | Unit-test style, assert per node | Must evaluate whole trajectories — see Agent Evaluation |
| Failure modes | A node errors; easy to locate | Drift, infinite loops, error accumulation; hard to locate |
The tell: if every branch is written in code and the model only works inside nodes, it's a workflow; if a while loop wraps an LLM and the loop's exit condition depends on the model's output, it's an agent.
Anthropic's advice from the original is worth memorizing: avoid agentic systems when you can; if a single LLM call plus retrieval solves it, don't build a workflow; if a workflow solves it, don't build an agent. Agentic systems trade latency and cost for task performance, and the trade only pays off when the task is open-ended enough and the path unpredictable enough. A huge share of failed "agent projects" in the industry failed not because the model was weak but because they wrapped a workflow-shaped problem in agent-shaped complexity.
3. Agent vs. Chatbot: a Conversational Shell vs. an Action Core
Chatbot: a system whose entire deliverable is the conversation. You send a message, it replies, interaction over. Even with a top LLM underneath, if its output stops at text replies and nothing outside the chat history changes, it's a chatbot.
- Typical examples: plain chat on the ChatGPT web app, support Q&A bots, most "AI assistant" bots in workplace messengers.
Agent differs by side effects: it changes the external world through tool calls — modifying code, sending email, placing orders, updating databases. Conversation is just one of its interfaces to humans, not its output.
A practical detection method: read the system's logs. If the log alternates only "user input → model output," it's a chatbot; if tool_call and tool_result entries appear interleaved and those calls have real external effects, there's agent in there. (Having tool calls doesn't mean having autonomy — a Q&A bot that only ever calls one fixed search tool is essentially a chatbot with retrieval; see the RAG section below.)
Their engineering challenges differ entirely too: a chatbot's hard problems are multi-turn dialogue management and intent recognition; an agent's are context engineering, tool interface design (Tools & MCP), and runaway protection (Agent Security).
The tell: after the system finishes, has the world been changed? Unchanged → chatbot. Changed in pursuit of the goal → agent.
4. Agent vs. Copilot: the Co-Driver and the Designated Driver
Copilot: AI embedded in a human's workflow, aiming to amplify the human. It completes, suggests, drafts — but final decision authority and execution authority stay with the person. The name is the definition: a copilot sits in the passenger seat; the steering wheel is in your hands.
- Typical examples: GitHub Copilot's inline code completion, Office Copilot's document polish, the "explain this code" panel in an IDE.
Agent is the designated driver: you name the destination, it drives the whole way, and you're only called back at critical checkpoints (payment, submitting a PR) to confirm — a design pattern called human-in-the-loop; see Human-in-the-Loop.
This boundary grew blurrier through 2025–2026 because "copilot products are all agentifying": GitHub Copilot shipped a coding-agent mode that completes issues independently, and Cursor's Composer/Agent mode modifies code across files autonomously (see the Cursor case). Whether a given feature is a copilot or an agent right now isn't determined by the product's name but by its default trust configuration:
- Every change needs human confirmation before it takes effect → copilot;
- Can execute many steps consecutively, with humans reviewing after the fact → agent.
The tell: look at the "radius of actions that take effect without confirmation." Radius zero (suggestions only) → copilot. Radius spanning multiple consecutive actions → agent. Both modes can coexist inside one product.
The engineering significance of this distinction is approval and rollback design: copilot architecture assumes the human is present in real time; agent architecture must assume they're not — hence logs, checkpoints, and rollback. See Observability.
5. Agent vs. RAG: RAG Is a Component, Not an Agent
This pair gets asked to death in interviews and answered wrong most often.
RAG (Retrieval-Augmented Generation): a technique for feeding external knowledge to an LLM — the user's question arrives, relevant document chunks are retrieved first, stuffed into the prompt, and the model generates an answer grounded in those chunks. The concept comes from Lewis et al.'s 2020 paper (arXiv 2005.11401, NeurIPS 2020), originally a model-training scheme of "parametric memory + non-parametric memory"; in engineering today it's shorthand for any "retrieve → assemble prompt → generate" pipeline. See the RAG page.
The key point: classic RAG is a straight-line pipeline with zero autonomous decisions. The "retrieve" step is a fixed action written by the developer, not something the model initiates after judging "I need more information." Therefore:
- RAG by itself is not an agent, just as "an engine is not a car";
- RAG can be one of an agent's tools — when the model decides in its loop that "this question needs the knowledge base" and initiates retrieval, RAG becomes one tool in the agent's toolbox. Anthropic calls this minimal unit the augmented LLM (an LLM augmented with retrieval, tools, and memory) — the building block of all agentic systems.
There are intermediates, of course — so-called agentic RAG: query rewriting, multi-route retrieval, re-querying when results aren't good enough. Once "whether to keep searching and how" becomes the model's decision, it starts sliding toward the agent side. Which proves the classification criterion isn't "does it retrieve?" — it's always "who owns the control flow."
The tell: a fixed three-step pipeline of "question → retrieve → generate" is RAG; a model that autonomously chooses "time to query the knowledge base" inside its loop is an agent that happens to use RAG as a tool.
Classic RAG (not an agent):
Question ──► Retrieval (fixed step) ──► Assemble prompt ──► LLM generates ──► Answer
Agentic RAG (an agent using RAG as a tool):
┌──────────────────────────────┐
▼ │
Goal ──► LLM decides ──► Calls the retrieval tool ──► Observes the result
│ "Enough — I can answer now"
▼
Answer6. Agent vs. Agentic AI: Engineering Usage vs. Academic Usage
Colloquially the two are interchangeable, but the academic literature has begun to distinguish them. Sapkota, Roumeliotis, and Karkee's May 2025 paper "AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges" (arXiv 2505.10468, 800+ citations) proposes a three-tier taxonomy:
| Tier | Core mechanism | Structure | Key traits |
|---|---|---|---|
| Generative AI | Prompt → LLM → output | Single model | Content generation, passive response |
| AI Agent | Prompt → tool calls → LLM → output | LLM + tools | Task-specific tool use |
| Agentic AI | Goal → multi-agent orchestration → output | Multi-agent system | Persistent memory, cross-agent collaboration, goal-level autonomy |
By this academic usage: "AI Agent" means a single, task-level tool user; "Agentic AI" means a system-level paradigm with multi-agent collaboration, persistent memory, and complex goal decomposition. In other words, in academic contexts "agentic AI" is nearly synonymous with what we'd call multi-agent systems plus long-horizon autonomy.
Engineering circles don't strictly honor the split — Anthropic uses "agentic systems" to cover both workflows and agents, a much broader sense. Practical advice: in papers, follow the author's definition; in documents, define before you use; in conversation, don't fight over it. But if you're writing it into a résumé or a technical proposal, the least attackable usage is: "agentic" as a system property ("this system has autonomy") and "agent" for concrete entities ("this system contains two agents").
The tell: "Agentic AI" in an academic paper title — assume it discusses multi-agent / goal-level-autonomous system paradigms. "Agentic" in a vendor blog — assume it's just a marketing adjective meaning "agent-related."
7. Agent vs. Automation / RPA: Deterministic Scripts vs. Probabilistic Decisions
Traditional automation / RPA (Robotic Process Automation): deterministic rules simulating a human operating software — click this button, fill this field, read from column 3 of the Excel. UiPath, Automation Anywhere, and their generation are the flagships. Its decision table is enumerated by a human in advance; on exceptions it errors out or escalates to a person.
The essential difference between an Agent and RPA isn't the vague notion of "intelligence level" — it's the direction of error attribution:
- RPA failure = the rules didn't cover this — the environment changed (the button moved, a popup appeared) and the script died;
- Agent failure = the model judged wrong — it can see the popup and understand the page changed, but it may still respond incorrectly.
| Dimension | RPA / traditional automation | LLM Agent |
|---|---|---|
| Source of decisions | Predefined rules, selectors, coordinates | The model's understanding of context |
| Tolerance to change | Extremely low; any UI change breaks it | Relatively high; understands interfaces semantically |
| Reproducibility | Fully reproducible | Probabilistic; trajectories never repeat |
| Cost per execution | Near zero | Burns tokens every run |
| Suited scenarios | High-frequency, fixed, compliance-heavy flows | Low-to-medium frequency, variable, judgment-requiring flows |
One notable 2025–2026 trend: RPA vendors are collectively "agentifying," embedding LLM decisions into their existing flow orchestration to form hybrids of "deterministic skeleton + pockets of intelligent decisions." That's actually the most pragmatic engineering route — compliance-critical steps run on fixed rules, comprehension-heavy steps (reading email, classifying documents) go to the model.
The tell: if you can draw a truth table for every branch of the flowchart, it's automation/RPA; if some nodes' outputs depend on the model's understanding of unstructured input, that part is already agent.
8. Single-Agent vs. Multi-Agent: First Ask Whether Division of Labor Is a Real Need
Single-agent system: one LLM loop, one context, one toolset. Multi-agent system: multiple agent instances, each with an independent context, coordinated by an orchestrator or peer-to-peer protocols. Architecture details in Multi-Agent Systems.
Industry's actual experience is far more sober than framework marketing. Anthropic's "How we built our multi-agent research system" reports that multi-agent architectures significantly outperformed single agents on their research tasks, while explicitly noting token consumption several times higher (and orders of magnitude worse for chat-like scenarios) — hence suitable only for "tasks valuable enough." Cognition's (the Devin team) widely circulated "Don't Build Multi-Agents" goes further: multi-agent architectures fail systematically on tasks requiring tightly shared context — sub-agents each make decisions, decisions contradict each other, and blame is hard to attribute.
The rule of thumb compresses to one sentence: multi-agent solves "a single context window can't hold it / a single role's prompt can't cover it," not "this looks sophisticated."
Signals that multi-agent is justified:
- The task splits naturally into parallel pieces whose subtasks don't need to share intermediate state frequently (typical: breadth-first research);
- Different subtasks need radically different toolsets and system prompts (one queries a database, another writes a report);
- A single agent's context is bursting and you need context isolation to compress tokens (see Context Engineering).
Signals against: subtasks need to see each other's intermediate results at every step — splitting then only introduces decision inconsistency; stick with a single agent plus good tools.
The tell: draw the system's context boundaries. One LLM loop, one main context → single agent. Multiple loops each maintaining their own context, requiring explicit message passing → multi-agent.
9. Agent vs. Harness / Scaffold: the Model Is Not the Agent
These terms went mainstream in engineering circles in the second half of 2025, and by May 2026 Hugging Face's official blog published a whole terminology guide, "Harness, Scaffold, and the AI Agent Terms Worth Getting Right," to untangle them — a measure of the confusion. Its core formula:
Agent = Model + Harness- Model: the LLM itself. Text in, text out; no memory between calls, no loop, no hands. It can "express" an intent to call a tool but cannot execute anything itself.
- Scaffold: the behavior-definition layer around the model — system prompt, tool descriptions, output-format conventions, context-management strategy. It shapes how the model "sees the world." The term comes from cognitive science's scaffolding (learning theory) and was already standard in evaluation circles.
- Harness: the execution layer around the model — the loop that actually invokes the model, parses tool calls, executes tools, feeds results back into the context, and decides when to stop. Claude Code's and Codex's official documentation call everything beyond the model the harness (or agentic harness).
This vocabulary explains a very important 2025–2026 industry phenomenon: why the same model performs wildly differently inside different products. Claude Code and some open-source alternative can call the exact same model yet produce worlds-apart results — the difference is entirely the harness: how the system prompt is written, how tool schemas are designed, how context is compacted, how errors are recovered. As model vendors get stronger, the application-layer moat migrates toward harness engineering.
The practical implication for engineers: when evaluating an agent product, don't just ask "which model does it use" — ask how the harness is designed. When does it stop the loop? How does it manage the context window? How does it recover from tool failures? Does it have a verifiable completion signal? These questions get closer to the essence of agent engineering than "how is the prompt written." See the Agent Loop and Agent System Anatomy.
The tell: discussing "model capability" → model. Discussing "prompt + tool definitions + context strategy" → scaffold. Discussing "the execution loop + tool runtime + stop conditions" → harness. The three together are the agent you actually experience.
Don't flunk this in an interview
"How is your agent different from calling the API directly?" — the correct answer lives entirely at the harness layer: loop control, context management, the tool runtime, permissions and guardrails, observability. A candidate who can only answer "our prompts are really good" will be read as someone who has never built a real agent system.
10. Summary: One Comparison Table
| Concept | One-line definition | Typical example | Tell (see X → it's Y) |
|---|---|---|---|
| Workflow | LLM and tools orchestrated through predefined code paths | Ticket-classification pipeline | Every branch written in code |
| Agent | LLM dynamically directs its own process to finish the task | Claude Code fixing a bug | The while loop's exit is the model's decision |
| Chatbot | Deliverable stops at conversational text | Support Q&A bot | Logs contain no tool calls with side effects |
| Copilot | An amplifier with the human present, confirming step by step | Inline code completion | Zero-radius of actions taking effect without confirmation |
| RAG | The fixed pipeline of retrieve → assemble prompt → generate | Enterprise knowledge-base QA | Retrieval is a hard-wired step, not a model decision |
| Agentic AI (academic) | The paradigm of multi-agent collaboration + persistent autonomy | Multi-agent research system | Paper-speak for the system-level paradigm |
| RPA / automation | Flow execution driven by deterministic rules | Invoice-entry script | Every branch has a truth table |
| Multi-agent | Multiple agents with independent contexts collaborating | Parallel research system | Explicit agent-to-agent message passing exists |
| Harness | The execution-and-control system beyond the model | Claude Code itself | The discussion covers loops, tool runtimes, stop conditions |
11. The Decision Flowchart: How to Choose for a New Requirement
This is where the whole article lands. Facing a new requirement, ask yourself four questions in the order below — the order itself is Anthropic's advice: start from the simplest solution; complexity is earned, not default.
A requirement comes in
│
Q1: Can a single LLM call solve it?
(translation, summarization, classification, rewriting…)
│ │
yes no
│ ▼
Just call the Q2: Is knowledge what's missing?
API. Don't (internal docs, live data,
over-engineer private domain facts)
│ │ │
│ yes no — what's missing is
│ │ "the ability to act"
│ ▼ │
│ Add RAG ▼
│ (retrieval + Q3: Can the execution
│ generation, steps be enumerated
│ fixed in advance?
│ pipeline) │ │
│ yes no — open-ended,
│ │ ever-changing path
│ ▼ │
│ Write a Workflow │
│ (code-orchestrated: │
│ predictable, cheap,│
│ easy to test) │
│ │ ▼
│ │ Use an Agent
│ │ (autonomous LLM loop
│ │ + tools + guardrails)
│ │ │
│ │ ▼
│ │ Q4: Does a single agent's
│ │ context overflow / can one
│ │ role not cover it all?
│ │ │ │
│ │ no yes
│ │ │ ▼
│ │ Stay single- Multi-agent
│ │ agent; tune system (and
│ │ tools & confirm the
│ │ context token budget)
▼ ▼
★ Shared prerequisite: whichever path you pick,
build an eval set and a baseline first (see /advanced/evaluation).
Otherwise you can't prove the more complex option is actually better.Three usage notes:
- Q2 and Q3 aren't mutually exclusive. Real systems are often hybrids: an agent's toolbox holds a RAG tool; a workflow node embeds a small agent. Classify by the backbone, compose locally.
- The answer to Q4 is "no" most of the time. As of 2026, a single agent + good tools + good context management covers the overwhelming majority of production needs. Multi-agent is the specialized answer for the few high-value, parallelizable, context-exploding tasks.
- Complexity is earned upward, never faked downward. Ship the workflow version first, collect failure cases, and use data to prove "fixed paths genuinely can't cover this" before adopting an agent — that process also builds your eval set along the way. Doing it in reverse (agent first, then simplify) almost always stalls at "we don't actually know where it's better or worse."
One sentence to remember it all
A workflow's path is drawn by a human; an agent's path is walked by the model; RAG is the component that feeds knowledge; a copilot keeps the steering wheel in your hands; the harness is the machinery beyond the model that actually does the work. To classify any system, ask one question: who made that decision?
References
- Building Effective Agents — Anthropic Engineering — the original source of the authoritative workflow-vs-agent distinction, including the five workflow patterns and selection advice
- A Practical Guide to Building Agents — OpenAI — OpenAI's productized definition of agents and the framework for judging when to build one
- Harness, Scaffold, and the AI Agent Terms Worth Getting Right — Hugging Face — the official agent-terminology guide published May 2026; source of the "Agent = Model + Harness" formula
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges — Sapkota et al.'s three-tier academic classification of Generative AI / AI Agent / Agentic AI (arXiv 2505.10468)
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al.'s original 2020 RAG paper (NeurIPS 2020)
- Claude Code Docs: Glossary — Claude Code's official definitions of agentic harness and related terms
- How we built our multi-agent research system — Anthropic — the first-hand engineering report on multi-agent gains and their multiples-of-token costs