Skip to content

LLM-Based Agents

At a glance LLM-based Agents upgrade large models from 'answering questions' to 'completing tasks': brain + tools + memory + planning loop. This article breaks down ReAct patterns, function calling and MCP, short-term/long-term memory, multi-agent collaboration, representative cases, and risk boundaries.

This page contains time-sensitive content, current as of 2025-06; job descriptions, rankings, product features, and other information may have changed. Please verify with original sources before citing.

LLM-Based Agents ​

LLM-based agents are systems that use large language models as the "brain" and autonomously call tools, use memory, and close the task loop through a "perceive → plan → act → observe" cycle. If ChatGPT taught large models to "speak," Agents teach them to "work" — booking flights, writing code, operating browsers, managing data. They are the most important evolution direction for large model applications post-2024, also known as the next stop from "chatting to working."

I. What Is an Agent: Four Components, All Required ​

ComponentRoleTypical Implementation
Brain (LLM)Understand goals, make decisions, generate action plansGPT-4o, Claude, DeepSeek — any strong model
Tools (Tools)Let the model interact with the external world (retrieval, code, APIs, browsers)Function calling, MCP, code interpreter, browser operation
Memory (Memory)Short-term: current task context; Long-term: cross-session experience and knowledgeConversation history, vector DB, file system
Loop (Loop)Repeatedly "think-act-observe" until the task is completeReAct, Plan-and-Execute, reflection

One sentence: Agent = LLM brain + tool limbs + memory inventory + loop-driven. Among these, "tools" are often a RAG retrieval system (see RAG: Retrieval-Augmented Generation) or external APIs.

One sentence to remember the essence of Agents: they bind "model judgment" with "program determinism" — if judgment is wrong, guardrails catch it; if program execution is wrong, logs hold accountability.

Beginners often confuse Agents with "long prompts." The difference: prompt engineering is "static" — instructions and examples are dumped into context once, model generates once; Agents are "dynamic" — the model makes repeated decisions in a loop, and each action changes the next input (tool results, memory updates). An operational litmus test: if a task can be completed in one generation, without tools and environment feedback, it's a prompt engineering problem; if a task needs multiple steps, depends on external state, and may fail mid-way and retry, it's an Agent problem. Of course they're on a continuum: a simple ReAct loop can also be seen as "multi-turn prompting with tools." This boundary also determines the learning order: get prompting right first (see Prompting), then introduce tools and loops.

Loop diagram:
while task not complete:
    Thought:   Decide the next step based on current state
    Action:    Call tool / generate content
    Observation: Read tool return / environment feedback
    Memory (Update): Write key info back to short/long-term memory

II. ReAct Pattern: Reasoning and Action Interleaved ​

ReAct (Reasoning + Acting) is the core operating mode for Agents, proposed by Yao et al. in 2022 (paper in references). Its essence is letting the model output "reasoning traces" and "tool calls" interleaved — think one step, do one step, observe one step:

Thought 1: I need to check Beijing's weather, call the weather tool first.
Action 1:  weather_search(city="Beijing")
Observation 1: {"temperature": 25, "weather": "sunny"}
Thought 2: Now I have the weather, generate clothing advice.
Action 2:  Answer user: "Beijing is 25°C, sunny today. Wear a light jacket."

ReAct's value: reasoning makes action well-grounded, observation makes reasoning feedback-rich. Compared to "generating an answer in one shot," ReAct significantly improves success rates on complex tasks requiring multi-step dependencies; the reasoning trace is also naturally auditable (every step has a reason why).

After ReAct, the research community advanced along two directions: "plan first, execute later" (Plan-and-Execute types), breaking long tasks into verifiable sub-plans to reduce the probability of getting lost in the loop; and "post-action reflection" (Reflexion types), letting the model rewrite its own action strategy based on failure feedback. The former suits structured tasks; the latter suits trial-and-error tasks (coding, solving math). Both can combine with ReAct into more robust loops; the tradeoff determines whether the Agent is "nervous at every step" or "steady and sure." More detailed comparison in Prompting Practice.

III. Tool Calling: Function Calling and MCP ​

1. Function Calling: Letting the Model "Command" Functions ​

In June 2023, OpenAI added function calling to its API: pass the JSON Schema of available functions to the model, and the model outputs structured JSON specifying "which function to call, with what parameters," which the host program executes and returns results to the model:

Available tools: [
  {name: "get_weather",  params: {city: string}},
  {name: "book_flight",  params: {from, to, date}}
]
Model output: {"name": "book_flight", "arguments": {"from":"Beijing","to":"Shanghai","date":"2025-06-01"}}

Function calling makes "model + arbitrary API" a standard capability: check inventory, send email, query databases, operate internal enterprise systems. Its engineering points (structured output, parameter validation, failure retry) are in Prompting Practice.

2. MCP: Standardizing Tools ​

In November 2024, Anthropic open-sourced Model Context Protocol (MCP), aiming to be the "USB-C port for AI apps": defining a unified protocol so any model framework can connect to file systems, databases, browsers, and dev tools in the same way. The MCP host-client-server model:

Model framework (Host) ←MCP protocol→ MCP Client ↔ MCP Server (exposing resources/tools/prompts)

MCP solves the fragmentation problem of the tool ecosystem: developers write one MCP server, and any Agent framework can reuse it. By 2025, MCP has become the de facto standard for open-source Agent tool integration.

The next question after toolification is: who guarantees "tool reliability"? Wrong tool Schema, missing parameter validation, and return format drift all cause the Agent's next decision to be built on wrong input. Engineering practice answers in three layers: first, "Schema is a contract" — tool descriptions must be precise to each parameter's meaning, type, and examples; second, "output validation" — call parameters generated by the model must pass validation before execution; third, "failure semantics" — tools should return structured error codes, not random text. Together, these three are the direct solution to the two pitfalls of "tool abuse" and "invalid calls" in Common Pitfalls and Anti-Patterns.

IV. Memory: Short-Term and Long-Term ​

Memory TypeCarrierLifespanUse
Short-term (working memory)Conversation history, context windowSingle taskMaintain current task state and goals
Long-term (knowledge/experience)Vector DB, database, file systemCross-sessionRemember user preferences, historical decisions, accumulate experience
Skill memoryPrompt templates, tool usageLong-termReuse proven problem-solving patterns

Short-term memory is limited by Context and Long Context, needing summarization/compression; long-term memory is usually implemented via vector retrieval (the RAG approach) — an Agent's long-term memory is essentially "using RAG as the Agent's memory store." The underlying retrieval and memory mechanisms are in RAG: Retrieval-Augmented Generation.

The most common Agent memory failure isn't "can't remember" but "misremember and confuse": short-term memory gets crowded by irrelevant info (context window filled with logs), long-term memory retrieves irrelevant old experience (vector retrieval drift), and memory updates overwrite important info. Engineering-wise, memory needs "write auditing" and "rollback" — every write has a source and timestamp, and can be reverted if needed. Treating memory as "a database needing version management" rather than "a conversation draft pad" is a must for production-grade Agents.

There's also a "compliance" dimension of memory often overlooked: long-term memory accumulates user personal information (preferences, health history, work content), and if leaked or misused, the risk far exceeds single-turn conversations. When designing memory systems, you must at least answer: what's stored, who can read it, how long it's retained, and how to delete it. Both the EU GDPR's "right to be forgotten" and China's Personal Information Protection Law directly constrain such scenarios. Designing memory as "a protected database" rather than "a conversation cache" is the baseline requirement for production Agents. Compliance frameworks in Safety and Risk.

V. Planning and Reflection ​

Relying solely on ReAct's "one step at a time" leads to getting lost in long tasks, so the industry has evolved more structured planning:

StrategyApproachSuitable For
CoT / ReActReasoning and action interleavedGeneral tasks (see Prompting)
Plan-and-ExecuteGenerate full plan first, then execute step by stepMulti-step, clear sub-tasks
ReflectionSelf-evaluate and correct after actionWriting, coding, iterative tasks
Tree search / multi-path samplingExplore multiple paths simultaneously, pick the bestHigh-value, high-cost-of-error

Letting Agents "think more" isn't free. Each reasoning step at the planning layer consumes tokens and time: an Agent with a full planning + reflection loop may consume 10–50× the tokens of a "direct answer." Engineering-wise, you need "task-level grading": simple questions go to "direct answer" (zero planning), medium tasks go to "one plan + execution," and only complex tasks get "planning + reflection + retry." The key to grading is the cost of error: writing an email has low error cost → direct answer; deleting a database has high error cost → plan first, review, then execute. This "cost-aware autonomy" design is an optimization point more immediately impactful than "stronger models" in Agent engineering. Related cost-control tips in Common Pitfalls and Anti-Patterns.

The essence of an Agent is "procedural reasoning"

The stronger the model's planning capability, the more reliable the Agent — which is why Agent capability collectively jumped when o1-type reasoning models appeared. Agent ceiling = model reasoning capability × tool quality × memory management.

VI. Multi-Agent Collaboration: Letting Agents Divide Labor ​

Multi-agent systems give different roles (PM, programmer, tester, reviewer) an Agent each, collaborating via messages to complete complex engineering tasks:

Product Manager Agent ──requirements──> Architect Agent ──design──> Developer Agent ──code──> Tester Agent
     ▲                                                              │
     └────────────────── defect feedback & iteration ───────────────┘
FrameworkCharacteristics
AutoGPTThe first autonomous Agent (2023.3), single Agent autonomously breaks down tasks
BabyAGITask queue + execution + memory, minimal playable prototype
MetaGPTMulti-role collaboration, simulating a software company SOP (2023.6)
ChatDevVirtual software company, dialogue-based pipeline development
OpenAI Swarm / Agent SDKLightweight multi-Agent orchestration
Commercial productsManus (general Agent), Devin (software engineer), Operator (browser operation)

Note: multi-Agents aren't a silver bullet — each additional Agent adds another layer of error propagation and cost; many tasks are "one Agent + good tools" and are more stable. Multi-Agent's value is in "clear role division, well-defined interfaces" scenarios (e.g., software engineering, content production pipelines).

Another "hidden cost" of multi-Agent systems is coordination overhead: roles need to pass state, synchronize cognition, and arbitrate conflicts. Experiments show that after Agent count exceeds 3–5, collaboration benefits are often offset by communication and error propagation. A more practical pattern is "layered": one "supervisor Agent" responsible for task decomposition and result acceptance, several "execution Agents" each completing sub-tasks, and optionally a "reviewer Agent." This "few and precise" hierarchical structure is far more stable than "flat-level free-for-all." For framework support, LangGraph's graph orchestration and OpenAI Agent SDK's hierarchical model both suit this design well (see Framework and Tool Selection).

VII. Representative Cases ​

CaseTimePositioningLesson
AutoGPT2023.3Autonomous task breakdown, open-source AgentIgnited the Agent concept, but poor stability and cost out of control, proving "loop alone isn't enough"
MetaGPT2023.6Multi-role software companySOP structuring reduces multi-Agent collaboration entropy
Devin2024.3AI software engineerLong tasks + coding toolchain + cloud environment, led SWE-bench at the time
OpenAI Operator2025.1Browser operation AgentConnected Agents to real webpages; paywalls/captchas become new barriers
Manus2025.3General autonomous AgentAsynchronous cloud execution, proving "task delegation" business model

Common pattern: successful Agents are the four-piece set of "model × environment × toolchain × evaluation" — changing only the brain (model) without engineering won't make the Agent stronger.

Representative cases give another lesson — the "pace of speed vs. stability": AutoGPT exploded the concept in the shortest time but faded quickly due to stability and cost; Devin and Manus narrowed task boundaries and execution environments, bringing Agents to "usable." Behind this is the Agent-specific "productization trap" — the "one-click complete" behind demo videos is often dozens of retries, manual intervention, and fixed environments. When evaluating an Agent product, don't just look at advertised completion rates — ask three questions: on what proportion of real inputs does it pass? When it fails, does it degrade gracefully or fail silently? What is the per-task cost? These three answers determine whether an Agent can go from "demo piece" to "productivity tool."

VIII. Risks: Prompt Injection, Loss of Control, and Cost ​

Agents have tools, permissions, and autonomous action — the risk is an order of magnitude higher than pure dialogue:

RiskManifestationMitigation
Prompt injectionMalicious instructions hidden in web/docs induce Agent to executeTool parameter validation, content/instruction isolation, least privilege
Tool abuseAgent accidentally calls high-cost/harmful APIsPermission whitelist, human-in-the-loop gate
Loss of control / unpredictabilityDrifts from goals in long loops, self-reinforces errorsStep limit, budget limit, checkpoint resume, sandbox isolation
Cost explosionLoop token consumption, API callsRouting (simple tasks go to small models), budget circuit breaker
Data leakageAgent sends private data to external toolsOutbound audit, desensitization, log retention

Prompt injection is the most unique Agent risk: attackers inject instructions through web/docs the Agent reads, turning the Agent into the attacker's tool. Defense details and security baselines are in Safety and Risk and Common Pitfalls and Anti-Patterns.

Don't give Agents a "master key"

An Agent's permissions should be "the minimum set needed to complete the task," with action logs recorded throughout. Any Agent that can autonomously call APIs should have: budget limit, step limit, sandbox, audit log — all four are non-negotiable.

IX. Agent System Engineering and Evaluation ​

1. Engineering Layers: An Agent Is More Than a Loop ​

LayerResponsibilityKey Question
Loop engineDrive "think-act-observe" until terminationTermination condition, step limit, retry strategy
Tool registrationExpose functions, APIs, MCP servers to the modelAccurate Schema, parameter validation, idempotency
State managementMaintain conversation history, memory, intermediate productsLong-task context compression, state snapshots
ObservabilityLog every step of reasoning and tool callsAudit log, checkpoint resume, playback

2. ReAct Prompt Template Example ​

Available tools:
- search(query): web search, returns webpage summary list
- calc(expr): execute math computation
- read_url(url): fetch webpage body

Output format: strictly follow these three steps in a loop:
Thought: <current judgment>
Action: <tool_name>(params)
Observation: <tool_return>

When no further tool call is needed, output:
Answer: <final answer>

More patterns for template engineering are in Prompting Practice.

3. Tool Definition JSON Schema ​

{
  "name": "book_flight",
  "description": "Book a flight (must validate date format and cabin class)",
  "parameters": {
    "type": "object",
    "properties": {
      "from": {"type": "string", "description": "Departure city"},
      "to":   {"type": "string", "description": "Arrival city"},
      "date": {"type": "string", "description": "Date YYYY-MM-DD"}
    },
    "required": ["from", "to", "date"]
  }
}

4. Agent Evaluation: Four Categories of Metrics, All Required ​

CategoryMetricsNotes
Task completionCompletion rate, success rate, quality scoreMeasured via real/simulated task sets
Tool usageCall accuracy, parameter error rate, invalid call rateExposes tool Schema quality issues
Cost efficiencyToken consumption, steps, API costLong loops can spiral; budgets are mandatory
SafetyJailbreak success rate, permission overreach, data leakageSee risks section above

5. State Management and Long-Task Engineering Details ​

The biggest gap between production-grade and demo-grade Agents is state management. Demo agents stuff all history into context — the longer the task, the more they "forget things" or "blow past context"; production agents need: ① structured task state (goals, sub-tasks, done/failed checklists) saved externally, not relying on model memory; ② context layering — high-frequency working memory in the window, low-frequency long-term memory in the vector DB (see RAG: Retrieval-Augmented Generation); ③ checkpoint resume — recover from the incomplete step after process crash; ④ idempotent tools — running the same tool call repeatedly doesn't produce side effects (the most common pitfall in Agent engineering). These engineering details determine whether an Agent can go from "demo works" to "production reliable." The more complete engineering checklist is in Common Pitfalls and Anti-Patterns.

Evaluation system setup is in Evaluations in Practice and Evaluation and Benchmarks.

5. Representative Framework Comparison ​

FrameworkCharacteristicsSuitable For
LangGraphGraph orchestration, fine-grained loop controlComplex workflows, enterprise-grade
AutoGenMulti-Agent dialogue collaborationResearch prototypes
OpenAI Agent SDKLightweight, official ecosystemQuick integration with GPT series
Claude Code / CursorTerminal/IDE AgentSoftware dev assistant
Dify / low-code platformsVisual orchestration + RAGBusiness users building apps

Complete framework selection methodology in Framework and Tool Selection; RAG as Agent memory store in RAG in Practice.

One main thread for Agent landing

From "single-turn Q&A" to "single Agent + tools" to "multi-Agent collaboration," each step requires re-evaluating ROI and loss-of-control risk. Most business value comes from the second step; the third is only worth it when division of labor is clear. Terminology reference in Glossary.

X. Boundary with Another Agent Handbook ​

The boundary between this handbook (The Large Model Handbook) and another series, the Agent Handbook:

TopicBelongs To
LLM brain: reasoning, prompting, alignment, hallucinationThe Large Model Handbook (this page and concepts pages)
Agent cognitive patterns: ReAct, planning, reflection, memory architectureThis page (case perspective) + Large Model Handbook concepts
Harness (runtime shell): execution loop engine, permission sandbox, lifecycle management, observability, concurrent schedulingAgent Handbook
Tool protocols (MCP), function calling specsThis page intro + Harness engineering implementation
Multi-Agent orchestration frameworks, deployment opsAgent Handbook

In short: this page answers "what the Agent's brain can do, how to make decisions"; the Agent Handbook answers "how the Agent's body executes reliably and safely." Readers building production-grade Agent systems should read this page together with the Agent Handbook.

A final recommendation on "when to read which": if you're doing prototype validation, evaluating "whether the model can be a brain," this page suffices; if you're running it as 7×24 production service (concurrency, permissions, observability, rollback), go read the Agent Handbook. Read both to understand both brain and body.

XI. Agent Landing and Future ​

1. Agent Commercialization Scenarios ​

ScenarioFormValueMaturity
Customer service / pre-salesRetrieval + ticketing + scripts7×24 responseHigh
Software developmentCode generation + self-test + fixMulti-fold efficiency improvementMed-high
Data and analyticsNatural language data query + reportsEveryone can analyzeMedium
Personal assistantSchedule, email, travel delegationSave transactional timeMedium
Enterprise process automationApproval, reconciliation, risk controlReplace RPA scenariosMed-low
Research assistantLiterature retrieval + experiment analysisAccelerate researchEarly

2. Prototype-to-Production Checklist ​

  • [ ] Clear task boundaries (what not to do, how far to go)
  • [ ] Tool whitelist and least privilege
  • [ ] Step limit, token budget, timeout circuit breaker
  • [ ] Full-chain logging and playback (traceable when issues arise)
  • [ ] Human approval gate (high-risk actions)
  • [ ] Evaluation set (task completion rate, cost, safety)
  • [ ] Gray release and rollback mechanism
  1. Reasoning model empowerment: o1-type "slow thinking" significantly improves Agent planning and error correction (see Frontier Progress);
  2. Protocol standardization: after MCP, standardized protocols for tools, identity, and permissions will continue to converge;
  3. From single-task to multi-task: Agents go from "completing one task" to "managing long-term goals";
  4. Multimodal perception: Agents that can read screens and understand meetings open new scenarios (see Multimodal LLMs).

4. Common Misconceptions and Anti-Patterns ​

MisconceptionCorrect Approach
Handing complex tasks entirely to AgentDecompose + human approval gate
Giving full permissions for "efficiency first"Least privilege + audit
Infinite loop with no limitStep/budget dual circuit breaker
Only testing capability, ignoring costInclude cost metrics in evaluation
Ignoring prompt injectionContent/instruction isolation, parameter validation

5. FAQ Quick Answers ​

QuestionQuick Answer
What's the relationship between Agent and RAG?RAG is one type of tool/memory for an Agent (see RAG: Retrieval-Augmented Generation)
Single Agent or multi-Agent?Most start with one Agent + tools; add multi-Agent only when division of labor is clear
Difference between function calling and MCP?FC is an interface capability; MCP is a standardized protocol
When do Agents lose control?No limit + high permissions + long loop — all three needed to avoid
How to evaluate Agents?The four-piece set: completion rate + tool accuracy + cost + safety
How does this page differ from the Agent Handbook?This page covers brain and cognition; Agent Handbook covers execution shell

One sentence to remember Agent landing

Agent value = task value × success rate − cost − risk. Pick tasks with "clear rules, fast feedback, reversible" first, solidify evaluation and guardrails, then gradually expand autonomy.

The ultimate measure of an Agent isn't "what it can do" but "what it does reliably" — write both success rate and cost into the acceptance criteria.

XII. Key Papers and Framework Resources for Agents ​

1. Key Papers at a Glance ​

Paper / WorkYearOne-Sentence Contribution
ReAct2022Founding paradigm of reasoning + acting interleaving
Toolformer2023Model self-teaches to use tools
HuggingGPT2023LLM-orchestrated multi-model tasks
Reflexion2023Language-feedback driven self-improvement
SWE-bench2023Real software engineering task benchmark
AgentBench2023Cross-environment Agent evaluation
MCP protocol2024Tool integration standardization

2. Evaluation Benchmarks ​

BenchmarkScenarioNotes
SWE-benchSoftware engineeringReal GitHub issue fixing
AgentBenchMulti-environmentOS, databases, web, etc.
GAIAGeneral assistantReal-world task sets
τ-benchTool callingTool use accuracy

3. Framework and Tool Resources ​

FrameworkCharacteristicsLearning Difficulty
LangGraphGraph orchestration, production-gradeMedium
AutoGenMulti-Agent dialogueLow
CrewAIRole-based multi-AgentLow
OpenAI Agent SDKOfficial lightweightLow
Claude CodeTerminal coding AgentLow
DifyVisual + RAG + AgentVery low

Action advice for beginners: don't rush into multi-Agents. Recommended entry path: ① use a ready API + function calling to build a "single Agent + three tools" minimal system (e.g., check weather, check calendar, send email); ② add a ReAct loop and logging, observe each step's decision quality; ③ introduce memory and evaluation sets, quantify completion rate and cost; ④ finally attempt multi-Agent and MCP. Each step has a clear, verifiable output, and exactly covers the knowledge loop from prompting to deployment. An Agent's capability comes from the design of loops and tools, not just model selection — start building, and cognition accelerates quickly.

XIII. Ethics and Responsibility for Agents ​

1. The Accountability Question ​

When Agents act autonomously, "who's responsible for the result" becomes blurry:

StageResponsible PartyRecommendation
Instruction understandingUserDefine task boundaries clearly
Tool callingDeveloperLeast privilege + audit
Final outputDeployerHuman approval gate + error correction
Data usageDeployerPrivacy and compliance assessment

2. Loss of Control and Misuse Risks ​

  • Unauthorized access: Agent accessing sensitive systems without authorization;
  • Prompt injection: external content inducing the Agent to execute malicious instructions;
  • Amplifying errors: wrong decisions automatically executed and spread (e.g., mass email, data deletion);
  • Fraud abuse: deepfakes, automated social engineering.

3. Principles for Responsible Deployment ​

  1. Default least privilege: always give the Agent less permission than the upper limit of need;
  2. Human approval for high-risk actions: deletion, payment, external publication must have human confirmation;
  3. Full traceability: every step of reasoning, tool call, and parameter is logged;
  4. Fail-safe default: stop on timeout/anomaly, don't continue;
  5. Continuous evaluation: include safety metrics in Agent evaluation sets (see Evaluations in Practice).

4. A Thought Experiment: When Not to Build an Agent ​

Responsibility also includes "restraint." When any of the following apply, prioritize not building an Agent: ① task failure has extremely high, irreversible cost (medical, financial decisions); ② task definition is so vague that even humans can't say what "success looks like"; ③ external dependencies are unstable (third-party APIs change anytime); ④ a stable traditional process already exists, and the Agent just "sounds cooler." An Agent's autonomy is a double-edged sword: it amplifies your execution capability, but also amplifies your mistakes. The prudent approach is "progressive autonomy" — manually review every step first, then delegate to partial steps, and only go fully autonomous after evaluation and guardrails are solid. This path echoes the three principles this page repeatedly emphasizes: "least privilege, fully auditable, fail-safe."

An Agent is not a "blame-shifting tool"

Handing a task to an Agent doesn't mean handing responsibility to the Agent. The deployer must be able to answer: what did it do, why did it do it, and who can stop it when things go wrong. If you can't answer these three questions, don't let the Agent face users or production systems directly.

A note: how to judge whether an Agent team is mature? Look at whether it treats "failure case review" as a fixed process — an Agent's progress comes from systematic attribution of failures, not from repeating successes.

XIV. Further Reading ​

References ​