Theme
LLM-Based Agents
LLM-based agents are systems that use large language models as the "brain" and autonomously call tools, use memory, and close the task loop through a "perceive → plan → act → observe" cycle. If ChatGPT taught large models to "speak," Agents teach them to "work" — booking flights, writing code, operating browsers, managing data. They are the most important evolution direction for large model applications post-2024, also known as the next stop from "chatting to working."
I. What Is an Agent: Four Components, All Required
| Component | Role | Typical Implementation |
|---|---|---|
| Brain (LLM) | Understand goals, make decisions, generate action plans | GPT-4o, Claude, DeepSeek — any strong model |
| Tools (Tools) | Let the model interact with the external world (retrieval, code, APIs, browsers) | Function calling, MCP, code interpreter, browser operation |
| Memory (Memory) | Short-term: current task context; Long-term: cross-session experience and knowledge | Conversation history, vector DB, file system |
| Loop (Loop) | Repeatedly "think-act-observe" until the task is complete | ReAct, Plan-and-Execute, reflection |
One sentence: Agent = LLM brain + tool limbs + memory inventory + loop-driven. Among these, "tools" are often a RAG retrieval system (see RAG: Retrieval-Augmented Generation) or external APIs.
One sentence to remember the essence of Agents: they bind "model judgment" with "program determinism" — if judgment is wrong, guardrails catch it; if program execution is wrong, logs hold accountability.
Beginners often confuse Agents with "long prompts." The difference: prompt engineering is "static" — instructions and examples are dumped into context once, model generates once; Agents are "dynamic" — the model makes repeated decisions in a loop, and each action changes the next input (tool results, memory updates). An operational litmus test: if a task can be completed in one generation, without tools and environment feedback, it's a prompt engineering problem; if a task needs multiple steps, depends on external state, and may fail mid-way and retry, it's an Agent problem. Of course they're on a continuum: a simple ReAct loop can also be seen as "multi-turn prompting with tools." This boundary also determines the learning order: get prompting right first (see Prompting), then introduce tools and loops.
Loop diagram:
while task not complete:
Thought: Decide the next step based on current state
Action: Call tool / generate content
Observation: Read tool return / environment feedback
Memory (Update): Write key info back to short/long-term memoryII. ReAct Pattern: Reasoning and Action Interleaved
ReAct (Reasoning + Acting) is the core operating mode for Agents, proposed by Yao et al. in 2022 (paper in references). Its essence is letting the model output "reasoning traces" and "tool calls" interleaved — think one step, do one step, observe one step:
Thought 1: I need to check Beijing's weather, call the weather tool first.
Action 1: weather_search(city="Beijing")
Observation 1: {"temperature": 25, "weather": "sunny"}
Thought 2: Now I have the weather, generate clothing advice.
Action 2: Answer user: "Beijing is 25°C, sunny today. Wear a light jacket."ReAct's value: reasoning makes action well-grounded, observation makes reasoning feedback-rich. Compared to "generating an answer in one shot," ReAct significantly improves success rates on complex tasks requiring multi-step dependencies; the reasoning trace is also naturally auditable (every step has a reason why).
After ReAct, the research community advanced along two directions: "plan first, execute later" (Plan-and-Execute types), breaking long tasks into verifiable sub-plans to reduce the probability of getting lost in the loop; and "post-action reflection" (Reflexion types), letting the model rewrite its own action strategy based on failure feedback. The former suits structured tasks; the latter suits trial-and-error tasks (coding, solving math). Both can combine with ReAct into more robust loops; the tradeoff determines whether the Agent is "nervous at every step" or "steady and sure." More detailed comparison in Prompting Practice.
III. Tool Calling: Function Calling and MCP
1. Function Calling: Letting the Model "Command" Functions
In June 2023, OpenAI added function calling to its API: pass the JSON Schema of available functions to the model, and the model outputs structured JSON specifying "which function to call, with what parameters," which the host program executes and returns results to the model:
Available tools: [
{name: "get_weather", params: {city: string}},
{name: "book_flight", params: {from, to, date}}
]
Model output: {"name": "book_flight", "arguments": {"from":"Beijing","to":"Shanghai","date":"2025-06-01"}}Function calling makes "model + arbitrary API" a standard capability: check inventory, send email, query databases, operate internal enterprise systems. Its engineering points (structured output, parameter validation, failure retry) are in Prompting Practice.
2. MCP: Standardizing Tools
In November 2024, Anthropic open-sourced Model Context Protocol (MCP), aiming to be the "USB-C port for AI apps": defining a unified protocol so any model framework can connect to file systems, databases, browsers, and dev tools in the same way. The MCP host-client-server model:
Model framework (Host) ←MCP protocol→ MCP Client ↔ MCP Server (exposing resources/tools/prompts)MCP solves the fragmentation problem of the tool ecosystem: developers write one MCP server, and any Agent framework can reuse it. By 2025, MCP has become the de facto standard for open-source Agent tool integration.
The next question after toolification is: who guarantees "tool reliability"? Wrong tool Schema, missing parameter validation, and return format drift all cause the Agent's next decision to be built on wrong input. Engineering practice answers in three layers: first, "Schema is a contract" — tool descriptions must be precise to each parameter's meaning, type, and examples; second, "output validation" — call parameters generated by the model must pass validation before execution; third, "failure semantics" — tools should return structured error codes, not random text. Together, these three are the direct solution to the two pitfalls of "tool abuse" and "invalid calls" in Common Pitfalls and Anti-Patterns.
IV. Memory: Short-Term and Long-Term
| Memory Type | Carrier | Lifespan | Use |
|---|---|---|---|
| Short-term (working memory) | Conversation history, context window | Single task | Maintain current task state and goals |
| Long-term (knowledge/experience) | Vector DB, database, file system | Cross-session | Remember user preferences, historical decisions, accumulate experience |
| Skill memory | Prompt templates, tool usage | Long-term | Reuse proven problem-solving patterns |
Short-term memory is limited by Context and Long Context, needing summarization/compression; long-term memory is usually implemented via vector retrieval (the RAG approach) — an Agent's long-term memory is essentially "using RAG as the Agent's memory store." The underlying retrieval and memory mechanisms are in RAG: Retrieval-Augmented Generation.
The most common Agent memory failure isn't "can't remember" but "misremember and confuse": short-term memory gets crowded by irrelevant info (context window filled with logs), long-term memory retrieves irrelevant old experience (vector retrieval drift), and memory updates overwrite important info. Engineering-wise, memory needs "write auditing" and "rollback" — every write has a source and timestamp, and can be reverted if needed. Treating memory as "a database needing version management" rather than "a conversation draft pad" is a must for production-grade Agents.
There's also a "compliance" dimension of memory often overlooked: long-term memory accumulates user personal information (preferences, health history, work content), and if leaked or misused, the risk far exceeds single-turn conversations. When designing memory systems, you must at least answer: what's stored, who can read it, how long it's retained, and how to delete it. Both the EU GDPR's "right to be forgotten" and China's Personal Information Protection Law directly constrain such scenarios. Designing memory as "a protected database" rather than "a conversation cache" is the baseline requirement for production Agents. Compliance frameworks in Safety and Risk.
V. Planning and Reflection
Relying solely on ReAct's "one step at a time" leads to getting lost in long tasks, so the industry has evolved more structured planning:
| Strategy | Approach | Suitable For |
|---|---|---|
| CoT / ReAct | Reasoning and action interleaved | General tasks (see Prompting) |
| Plan-and-Execute | Generate full plan first, then execute step by step | Multi-step, clear sub-tasks |
| Reflection | Self-evaluate and correct after action | Writing, coding, iterative tasks |
| Tree search / multi-path sampling | Explore multiple paths simultaneously, pick the best | High-value, high-cost-of-error |
Letting Agents "think more" isn't free. Each reasoning step at the planning layer consumes tokens and time: an Agent with a full planning + reflection loop may consume 10–50× the tokens of a "direct answer." Engineering-wise, you need "task-level grading": simple questions go to "direct answer" (zero planning), medium tasks go to "one plan + execution," and only complex tasks get "planning + reflection + retry." The key to grading is the cost of error: writing an email has low error cost → direct answer; deleting a database has high error cost → plan first, review, then execute. This "cost-aware autonomy" design is an optimization point more immediately impactful than "stronger models" in Agent engineering. Related cost-control tips in Common Pitfalls and Anti-Patterns.
The essence of an Agent is "procedural reasoning"
The stronger the model's planning capability, the more reliable the Agent — which is why Agent capability collectively jumped when o1-type reasoning models appeared. Agent ceiling = model reasoning capability × tool quality × memory management.
VI. Multi-Agent Collaboration: Letting Agents Divide Labor
Multi-agent systems give different roles (PM, programmer, tester, reviewer) an Agent each, collaborating via messages to complete complex engineering tasks:
Product Manager Agent ──requirements──> Architect Agent ──design──> Developer Agent ──code──> Tester Agent
▲ │
└────────────────── defect feedback & iteration ───────────────┘| Framework | Characteristics |
|---|---|
| AutoGPT | The first autonomous Agent (2023.3), single Agent autonomously breaks down tasks |
| BabyAGI | Task queue + execution + memory, minimal playable prototype |
| MetaGPT | Multi-role collaboration, simulating a software company SOP (2023.6) |
| ChatDev | Virtual software company, dialogue-based pipeline development |
| OpenAI Swarm / Agent SDK | Lightweight multi-Agent orchestration |
| Commercial products | Manus (general Agent), Devin (software engineer), Operator (browser operation) |
Note: multi-Agents aren't a silver bullet — each additional Agent adds another layer of error propagation and cost; many tasks are "one Agent + good tools" and are more stable. Multi-Agent's value is in "clear role division, well-defined interfaces" scenarios (e.g., software engineering, content production pipelines).
Another "hidden cost" of multi-Agent systems is coordination overhead: roles need to pass state, synchronize cognition, and arbitrate conflicts. Experiments show that after Agent count exceeds 3–5, collaboration benefits are often offset by communication and error propagation. A more practical pattern is "layered": one "supervisor Agent" responsible for task decomposition and result acceptance, several "execution Agents" each completing sub-tasks, and optionally a "reviewer Agent." This "few and precise" hierarchical structure is far more stable than "flat-level free-for-all." For framework support, LangGraph's graph orchestration and OpenAI Agent SDK's hierarchical model both suit this design well (see Framework and Tool Selection).
VII. Representative Cases
| Case | Time | Positioning | Lesson |
|---|---|---|---|
| AutoGPT | 2023.3 | Autonomous task breakdown, open-source Agent | Ignited the Agent concept, but poor stability and cost out of control, proving "loop alone isn't enough" |
| MetaGPT | 2023.6 | Multi-role software company | SOP structuring reduces multi-Agent collaboration entropy |
| Devin | 2024.3 | AI software engineer | Long tasks + coding toolchain + cloud environment, led SWE-bench at the time |
| OpenAI Operator | 2025.1 | Browser operation Agent | Connected Agents to real webpages; paywalls/captchas become new barriers |
| Manus | 2025.3 | General autonomous Agent | Asynchronous cloud execution, proving "task delegation" business model |
Common pattern: successful Agents are the four-piece set of "model × environment × toolchain × evaluation" — changing only the brain (model) without engineering won't make the Agent stronger.
Representative cases give another lesson — the "pace of speed vs. stability": AutoGPT exploded the concept in the shortest time but faded quickly due to stability and cost; Devin and Manus narrowed task boundaries and execution environments, bringing Agents to "usable." Behind this is the Agent-specific "productization trap" — the "one-click complete" behind demo videos is often dozens of retries, manual intervention, and fixed environments. When evaluating an Agent product, don't just look at advertised completion rates — ask three questions: on what proportion of real inputs does it pass? When it fails, does it degrade gracefully or fail silently? What is the per-task cost? These three answers determine whether an Agent can go from "demo piece" to "productivity tool."
VIII. Risks: Prompt Injection, Loss of Control, and Cost
Agents have tools, permissions, and autonomous action — the risk is an order of magnitude higher than pure dialogue:
| Risk | Manifestation | Mitigation |
|---|---|---|
| Prompt injection | Malicious instructions hidden in web/docs induce Agent to execute | Tool parameter validation, content/instruction isolation, least privilege |
| Tool abuse | Agent accidentally calls high-cost/harmful APIs | Permission whitelist, human-in-the-loop gate |
| Loss of control / unpredictability | Drifts from goals in long loops, self-reinforces errors | Step limit, budget limit, checkpoint resume, sandbox isolation |
| Cost explosion | Loop token consumption, API calls | Routing (simple tasks go to small models), budget circuit breaker |
| Data leakage | Agent sends private data to external tools | Outbound audit, desensitization, log retention |
Prompt injection is the most unique Agent risk: attackers inject instructions through web/docs the Agent reads, turning the Agent into the attacker's tool. Defense details and security baselines are in Safety and Risk and Common Pitfalls and Anti-Patterns.
Don't give Agents a "master key"
An Agent's permissions should be "the minimum set needed to complete the task," with action logs recorded throughout. Any Agent that can autonomously call APIs should have: budget limit, step limit, sandbox, audit log — all four are non-negotiable.
IX. Agent System Engineering and Evaluation
1. Engineering Layers: An Agent Is More Than a Loop
| Layer | Responsibility | Key Question |
|---|---|---|
| Loop engine | Drive "think-act-observe" until termination | Termination condition, step limit, retry strategy |
| Tool registration | Expose functions, APIs, MCP servers to the model | Accurate Schema, parameter validation, idempotency |
| State management | Maintain conversation history, memory, intermediate products | Long-task context compression, state snapshots |
| Observability | Log every step of reasoning and tool calls | Audit log, checkpoint resume, playback |
2. ReAct Prompt Template Example
Available tools:
- search(query): web search, returns webpage summary list
- calc(expr): execute math computation
- read_url(url): fetch webpage body
Output format: strictly follow these three steps in a loop:
Thought: <current judgment>
Action: <tool_name>(params)
Observation: <tool_return>
When no further tool call is needed, output:
Answer: <final answer>More patterns for template engineering are in Prompting Practice.
3. Tool Definition JSON Schema
{
"name": "book_flight",
"description": "Book a flight (must validate date format and cabin class)",
"parameters": {
"type": "object",
"properties": {
"from": {"type": "string", "description": "Departure city"},
"to": {"type": "string", "description": "Arrival city"},
"date": {"type": "string", "description": "Date YYYY-MM-DD"}
},
"required": ["from", "to", "date"]
}
}4. Agent Evaluation: Four Categories of Metrics, All Required
| Category | Metrics | Notes |
|---|---|---|
| Task completion | Completion rate, success rate, quality score | Measured via real/simulated task sets |
| Tool usage | Call accuracy, parameter error rate, invalid call rate | Exposes tool Schema quality issues |
| Cost efficiency | Token consumption, steps, API cost | Long loops can spiral; budgets are mandatory |
| Safety | Jailbreak success rate, permission overreach, data leakage | See risks section above |
5. State Management and Long-Task Engineering Details
The biggest gap between production-grade and demo-grade Agents is state management. Demo agents stuff all history into context — the longer the task, the more they "forget things" or "blow past context"; production agents need: ① structured task state (goals, sub-tasks, done/failed checklists) saved externally, not relying on model memory; ② context layering — high-frequency working memory in the window, low-frequency long-term memory in the vector DB (see RAG: Retrieval-Augmented Generation); ③ checkpoint resume — recover from the incomplete step after process crash; ④ idempotent tools — running the same tool call repeatedly doesn't produce side effects (the most common pitfall in Agent engineering). These engineering details determine whether an Agent can go from "demo works" to "production reliable." The more complete engineering checklist is in Common Pitfalls and Anti-Patterns.
Evaluation system setup is in Evaluations in Practice and Evaluation and Benchmarks.
5. Representative Framework Comparison
| Framework | Characteristics | Suitable For |
|---|---|---|
| LangGraph | Graph orchestration, fine-grained loop control | Complex workflows, enterprise-grade |
| AutoGen | Multi-Agent dialogue collaboration | Research prototypes |
| OpenAI Agent SDK | Lightweight, official ecosystem | Quick integration with GPT series |
| Claude Code / Cursor | Terminal/IDE Agent | Software dev assistant |
| Dify / low-code platforms | Visual orchestration + RAG | Business users building apps |
Complete framework selection methodology in Framework and Tool Selection; RAG as Agent memory store in RAG in Practice.
One main thread for Agent landing
From "single-turn Q&A" to "single Agent + tools" to "multi-Agent collaboration," each step requires re-evaluating ROI and loss-of-control risk. Most business value comes from the second step; the third is only worth it when division of labor is clear. Terminology reference in Glossary.
X. Boundary with Another Agent Handbook
The boundary between this handbook (The Large Model Handbook) and another series, the Agent Handbook:
| Topic | Belongs To |
|---|---|
| LLM brain: reasoning, prompting, alignment, hallucination | The Large Model Handbook (this page and concepts pages) |
| Agent cognitive patterns: ReAct, planning, reflection, memory architecture | This page (case perspective) + Large Model Handbook concepts |
| Harness (runtime shell): execution loop engine, permission sandbox, lifecycle management, observability, concurrent scheduling | Agent Handbook |
| Tool protocols (MCP), function calling specs | This page intro + Harness engineering implementation |
| Multi-Agent orchestration frameworks, deployment ops | Agent Handbook |
In short: this page answers "what the Agent's brain can do, how to make decisions"; the Agent Handbook answers "how the Agent's body executes reliably and safely." Readers building production-grade Agent systems should read this page together with the Agent Handbook.
A final recommendation on "when to read which": if you're doing prototype validation, evaluating "whether the model can be a brain," this page suffices; if you're running it as 7×24 production service (concurrency, permissions, observability, rollback), go read the Agent Handbook. Read both to understand both brain and body.
XI. Agent Landing and Future
1. Agent Commercialization Scenarios
| Scenario | Form | Value | Maturity |
|---|---|---|---|
| Customer service / pre-sales | Retrieval + ticketing + scripts | 7×24 response | High |
| Software development | Code generation + self-test + fix | Multi-fold efficiency improvement | Med-high |
| Data and analytics | Natural language data query + reports | Everyone can analyze | Medium |
| Personal assistant | Schedule, email, travel delegation | Save transactional time | Medium |
| Enterprise process automation | Approval, reconciliation, risk control | Replace RPA scenarios | Med-low |
| Research assistant | Literature retrieval + experiment analysis | Accelerate research | Early |
2. Prototype-to-Production Checklist
- [ ] Clear task boundaries (what not to do, how far to go)
- [ ] Tool whitelist and least privilege
- [ ] Step limit, token budget, timeout circuit breaker
- [ ] Full-chain logging and playback (traceable when issues arise)
- [ ] Human approval gate (high-risk actions)
- [ ] Evaluation set (task completion rate, cost, safety)
- [ ] Gray release and rollback mechanism
3. Future Trends for Agents
- Reasoning model empowerment: o1-type "slow thinking" significantly improves Agent planning and error correction (see Frontier Progress);
- Protocol standardization: after MCP, standardized protocols for tools, identity, and permissions will continue to converge;
- From single-task to multi-task: Agents go from "completing one task" to "managing long-term goals";
- Multimodal perception: Agents that can read screens and understand meetings open new scenarios (see Multimodal LLMs).
4. Common Misconceptions and Anti-Patterns
| Misconception | Correct Approach |
|---|---|
| Handing complex tasks entirely to Agent | Decompose + human approval gate |
| Giving full permissions for "efficiency first" | Least privilege + audit |
| Infinite loop with no limit | Step/budget dual circuit breaker |
| Only testing capability, ignoring cost | Include cost metrics in evaluation |
| Ignoring prompt injection | Content/instruction isolation, parameter validation |
5. FAQ Quick Answers
| Question | Quick Answer |
|---|---|
| What's the relationship between Agent and RAG? | RAG is one type of tool/memory for an Agent (see RAG: Retrieval-Augmented Generation) |
| Single Agent or multi-Agent? | Most start with one Agent + tools; add multi-Agent only when division of labor is clear |
| Difference between function calling and MCP? | FC is an interface capability; MCP is a standardized protocol |
| When do Agents lose control? | No limit + high permissions + long loop — all three needed to avoid |
| How to evaluate Agents? | The four-piece set: completion rate + tool accuracy + cost + safety |
| How does this page differ from the Agent Handbook? | This page covers brain and cognition; Agent Handbook covers execution shell |
One sentence to remember Agent landing
Agent value = task value × success rate − cost − risk. Pick tasks with "clear rules, fast feedback, reversible" first, solidify evaluation and guardrails, then gradually expand autonomy.
The ultimate measure of an Agent isn't "what it can do" but "what it does reliably" — write both success rate and cost into the acceptance criteria.
XII. Key Papers and Framework Resources for Agents
1. Key Papers at a Glance
| Paper / Work | Year | One-Sentence Contribution |
|---|---|---|
| ReAct | 2022 | Founding paradigm of reasoning + acting interleaving |
| Toolformer | 2023 | Model self-teaches to use tools |
| HuggingGPT | 2023 | LLM-orchestrated multi-model tasks |
| Reflexion | 2023 | Language-feedback driven self-improvement |
| SWE-bench | 2023 | Real software engineering task benchmark |
| AgentBench | 2023 | Cross-environment Agent evaluation |
| MCP protocol | 2024 | Tool integration standardization |
2. Evaluation Benchmarks
| Benchmark | Scenario | Notes |
|---|---|---|
| SWE-bench | Software engineering | Real GitHub issue fixing |
| AgentBench | Multi-environment | OS, databases, web, etc. |
| GAIA | General assistant | Real-world task sets |
| τ-bench | Tool calling | Tool use accuracy |
3. Framework and Tool Resources
| Framework | Characteristics | Learning Difficulty |
|---|---|---|
| LangGraph | Graph orchestration, production-grade | Medium |
| AutoGen | Multi-Agent dialogue | Low |
| CrewAI | Role-based multi-Agent | Low |
| OpenAI Agent SDK | Official lightweight | Low |
| Claude Code | Terminal coding Agent | Low |
| Dify | Visual + RAG + Agent | Very low |
Action advice for beginners: don't rush into multi-Agents. Recommended entry path: ① use a ready API + function calling to build a "single Agent + three tools" minimal system (e.g., check weather, check calendar, send email); ② add a ReAct loop and logging, observe each step's decision quality; ③ introduce memory and evaluation sets, quantify completion rate and cost; ④ finally attempt multi-Agent and MCP. Each step has a clear, verifiable output, and exactly covers the knowledge loop from prompting to deployment. An Agent's capability comes from the design of loops and tools, not just model selection — start building, and cognition accelerates quickly.
XIII. Ethics and Responsibility for Agents
1. The Accountability Question
When Agents act autonomously, "who's responsible for the result" becomes blurry:
| Stage | Responsible Party | Recommendation |
|---|---|---|
| Instruction understanding | User | Define task boundaries clearly |
| Tool calling | Developer | Least privilege + audit |
| Final output | Deployer | Human approval gate + error correction |
| Data usage | Deployer | Privacy and compliance assessment |
2. Loss of Control and Misuse Risks
- Unauthorized access: Agent accessing sensitive systems without authorization;
- Prompt injection: external content inducing the Agent to execute malicious instructions;
- Amplifying errors: wrong decisions automatically executed and spread (e.g., mass email, data deletion);
- Fraud abuse: deepfakes, automated social engineering.
3. Principles for Responsible Deployment
- Default least privilege: always give the Agent less permission than the upper limit of need;
- Human approval for high-risk actions: deletion, payment, external publication must have human confirmation;
- Full traceability: every step of reasoning, tool call, and parameter is logged;
- Fail-safe default: stop on timeout/anomaly, don't continue;
- Continuous evaluation: include safety metrics in Agent evaluation sets (see Evaluations in Practice).
4. A Thought Experiment: When Not to Build an Agent
Responsibility also includes "restraint." When any of the following apply, prioritize not building an Agent: ① task failure has extremely high, irreversible cost (medical, financial decisions); ② task definition is so vague that even humans can't say what "success looks like"; ③ external dependencies are unstable (third-party APIs change anytime); ④ a stable traditional process already exists, and the Agent just "sounds cooler." An Agent's autonomy is a double-edged sword: it amplifies your execution capability, but also amplifies your mistakes. The prudent approach is "progressive autonomy" — manually review every step first, then delegate to partial steps, and only go fully autonomous after evaluation and guardrails are solid. This path echoes the three principles this page repeatedly emphasizes: "least privilege, fully auditable, fail-safe."
An Agent is not a "blame-shifting tool"
Handing a task to an Agent doesn't mean handing responsibility to the Agent. The deployer must be able to answer: what did it do, why did it do it, and who can stop it when things go wrong. If you can't answer these three questions, don't let the Agent face users or production systems directly.
A note: how to judge whether an Agent team is mature? Look at whether it treats "failure case review" as a fixed process — an Agent's progress comes from systematic attribution of failures, not from repeating successes.
XIV. Further Reading
- Prompting — the prompting foundation for Agent reasoning and planning
- RAG: Retrieval-Augmented Generation — Agent long-term memory and knowledge tools
- Safety and Risk — prompt injection and Agent security
- Prompting Practice — function calling and structured output engineering
- RAG in Practice — Agent memory store building
- Common Pitfalls and Anti-Patterns — Agent loss of control and cost traps
- Frontier Progress — Agent and tool use research trends
References
- Yao et al. ReAct: Synergizing Reasoning and Acting in Language Models (2022) — ReAct paper (arXiv)
- Schick et al. Toolformer: Language Models Can Teach Themselves to Use Tools (2023) — Toolformer paper (arXiv)
- OpenAI. Function calling and other API updates (2023.6) — function calling official release
- Anthropic. Model Context Protocol official docs — MCP protocol documentation
- Significant-Gravitas. AutoGPT open-source repo — AutoGPT codebase
- MetaGPT open-source repo — MetaGPT codebase
- Cognition. Introducing Devin (2024.3) — Devin official release
- OWASP. Top 10 for Large Language Model Applications — LLM application security risk list