Appearance
Classic Papers, Deep-Dived
This page has the highest academic density on the whole site. We selected 10 papers that genuinely shaped the LLM Agent field and deep-dive into each in chronological order. The selection criterion is just one: open the source code of any Agent framework today, and you'll find its shadow.
Each paper follows a fixed template: publication info → the problem it tackles → core method → key experiments and data → limitations → influence on later systems. Short on time? Prioritize ReAct, Chain-of-Thought, and Reflexion — they come up in interviews most often (see the interview question bank).
How to read a paper
On the first pass, read only the abstract, Figure 1, and the main results table, answering three questions: what did it solve, how, and how well? Save the method details for the second pass. This page does the first pass for you; for the second, go back to the original.
0. Overview
| Paper | Year | One-line contribution | Later influence |
|---|---|---|---|
| Chain-of-Thought | 2022 | Gets the model to "write out its thinking," unlocking multi-step reasoning | The starting point of every reasoning paradigm; the Agent's Thought field |
| MRKL | 2022 | First systematic proposal of the "LLM + router + expert modules" architecture | LangChain's design blueprint; the origin of tool-routing ideas |
| ReAct | 2022 | Interleaves reasoning and action (Thought/Action/Observation loop) | The de facto Agent Loop standard, inherited by every framework |
| Toolformer | 2023 | The model teaches itself when and how to call APIs | The forerunner of OpenAI function calling; the tool-training paradigm |
| Reflexion | 2023 | Uses language feedback instead of gradient updates to achieve "reinforcement learning" | The stock component of self-reflection / self-correction mechanisms |
| Generative Agents | 2023 | A complete Agent architecture of memory stream + reflection + planning | The foundational work of memory-system design; multi-agent simulation |
| Tree of Thoughts | 2023 | Upgrades single-chain reasoning to backtrackable tree search | The pioneer of test-time search |
| Voyager | 2023 | Skill library + automatic curriculum enabling lifelong learning | The "executable skill library" idea influenced code agents |
| CodeAct | 2024 | Unifies the Agent action space with executable Python code | The default action paradigm of OpenHands and smolagents |
| SWE-agent | 2024 | Establishes the design discipline of the Agent-Computer Interface (ACI) | Tool interface design became a field of study in its own right |
You can group them into four layers: reasoning (CoT, ToT), action (MRKL, ReAct, Toolformer, CodeAct), learning and memory (Reflexion, Generative Agents, Voyager), and interface (SWE-agent). This maps one-to-one onto the component breakdown in Anatomy of an Agent.
1. Chain-of-Thought Prompting (2022)
Publication info: Wei et al., Google Research. arXiv:2201.11903 (submitted January 2022), published at NeurIPS 2022.
The problem it tackles
With standard prompting, the model maps straight from "question → answer." For arithmetic, commonsense, and symbolic reasoning tasks that require multi-step derivation, accuracy stays low even when the model is scaled to hundreds of billions of parameters. Fine-tuning helps, but annotating massive reasoning traces for every reasoning task isn't realistic.
Core method
Almost too simple for a milestone paper: in the few-shot examples, write out the "intermediate reasoning steps," and the model imitates that format, producing a reasoning chain before the answer.
# Standard prompting
Q: The cafeteria had 23 apples. It used 20 to make lunch and bought 6 more.
How many are left?
A: 9 ← the model just guesses; error-prone
# Chain-of-Thought prompting
Q: The cafeteria had 23 apples. It used 20 to make lunch and bought 6 more.
How many are left?
A: The cafeteria started with 23 apples. After using 20, 23 - 20 = 3 remain.
Then it bought 6 more, so now there are 3 + 6 = 9. The answer is 9.A zero-shot variant that later saw wide use appends "Let's think step by step" to the question, which similarly elicits a reasoning chain.
Key experiments and data
- On GSM8K math word problems, PaLM 540B + CoT reached about 57% accuracy, beating the previous best of fine-tuned GPT-3 with a verifier — without any fine-tuning.
- CoT is an emergent capability: below roughly 100B parameters it's nearly useless or even harmful — one of the most famous pieces of evidence that "scale unlocks capability."
- The gains appear consistently across arithmetic (GSM8K), commonsense (CSQA), and symbolic reasoning (Last Letter Concatenation), among other task types.
Limitations
- No error-correction mechanism once the chain goes wrong: one wrong step, every following step is wrong (error propagation).
- The reasoning is "generated text," not guaranteed to be faithful to the model's actual computation — later interpretability research repeatedly made this point.
- Pure text reasoning can't access external information, so hallucination still happens (exactly the gap ReAct set out to fill).
Influence on later systems
CoT is the cognitive core of the Agent. Every "thought" you see in a LangGraph trace, every "analyze first, then act" monologue in Claude Code's output, is a descendant of CoT. After 2024, reasoning models like OpenAI's o-series and DeepSeek-R1 essentially internalized CoT from a "prompt trick" into a "training objective." See Prompt Engineering and A Brief History.
2. MRKL Systems (2022)
Publication info: Karpas et al., AI21 Labs. arXiv:2205.00445 (May 2022), arXiv preprint (never went through a major conference, but its influence rivals any conference paper).
The problem it tackles
LMs in 2022 had three hard flaws: knowledge frozen at the training cutoff, unreliable arithmetic, and no access to private data. Rather than waiting for an omniscient giant model, AI21's answer was: acknowledge the limits of a single model and route around them with system architecture.
Core method
An MRKL (Modular Reasoning, Knowledge and Language) system = a router + a set of "expert modules":
User input
│
▼
┌──────────┐
│ Router │ ← can itself be an LM
└──────────┘
│ │ │
▼ ▼ ▼
┌────────────┐┌────────────┐┌────────────┐
│ Calculator ││ Search ││ Database │ ← discrete experts (symbolic, reliable)
└────────────┘└────────────┘└────────────┘
│ │ │
▼ ▼ ▼
Results returned to the LM to compose a natural-language answerExpert modules can be neural (another LM) or symbolic (a calculator, a SQL engine, a search engine). When routing is correct, symbolic experts guarantee exact, verifiable results.
Key experiments and data
The paper ran routing experiments on AI21's own Jurassic-1. The most instructive part isn't the success but the failure: on arithmetic tasks, when trained only on single-operation samples and then tested on two-step compound operations, some combinations collapsed outright (certain add-then-multiply and subtract-then-multiply combos performed near randomly). Reliability of a single tool does not imply reliability of composed calls — the earliest empirical evidence of the Agent composition-generalization problem.
Limitations
- The router is a new bottleneck: misroute, and the calculator will "faithfully compute the wrong thing."
- Every new expert needs training data or an adaptation layer — scaling costs are nontrivial.
- The ecosystem wasn't mature when the paper came out; it reads more like an architecture manifesto than a usable system.
Influence on later systems
MRKL named the whole idea of "LLM orchestrating external tools." LangChain's early docs cited it directly, and LangChain's Agent abstraction (the LLM decides which Tool to call) is MRKL in popular form. In today's MCP ecosystem, tool routing has exactly the same problem structure as in 2022. See Tools and MCP.
3. ReAct (2022)
Publication info: Yao et al., Princeton University & Google Research. arXiv:2210.03629 (October 2022), published at ICLR 2023.
The problem it tackles
CoT can "think" but can't "act" — its reasoning chain can't fetch external information, so factual errors run unchecked to the end; earlier action-based methods (directly generating action sequences) can't "think," so they can't adjust the plan when things fail. Reasoning and action were being studied as two separate problems.
Core method
Alternate the reasoning trace (Thought) with actions (Action), letting the model think and act within the same generation stream:
Question: The region where the eastern sector of the Colorado orogeny extends,
what is its elevation?
Thought 1: I need to search for "Colorado orogeny," find the region its eastern
sector extends to, then look up that region's elevation.
Action 1: Search[Colorado orogeny]
Observation 1: The Colorado orogeny was an orogeny in Colorado and surrounding
areas...
Thought 2: The passage doesn't mention the eastern sector's region. I need to
look up "eastern sector" further.
Action 2: Lookup[eastern sector]
Observation 2: The eastern sector of the orogeny extends to the High Plains...
Thought 3: The eastern sector extends to the High Plains. Now look up the
elevation of the High Plains.
Action 3: Search[High Plains elevation]
Observation 3: The High Plains rise from about 600 meters in the southeast to
the eastern border...
Thought 4: I have the answer.
Action 4: Finish[roughly 600–2000 meters]This loop is the prototype of every Agent Loop today. Thought lets the model decompose the task, track progress, and recover from anomalous Observations; Action lets it ground in the real world and fight hallucination.
Key experiments and data
- Knowledge-intensive tasks (HotpotQA, FEVER): ReAct, interacting with a minimal Wikipedia API, mitigated CoT's hallucination and error propagation.
- Interactive decision tasks: on ALFWorld (a text-based household environment) it beat imitation-learning/RL baselines by 34 percentage points in absolute success rate, and on WebShop (simulated online shopping) by 10 points — using only 1-2 in-context examples, with no training.
- Human evaluation showed ReAct's trajectories were more interpretable than no-reasoning baselines.
Limitations
- Heavily dependent on example quality in the prompt; performance drops fast when examples don't match the task.
- Once a trajectory starts down the wrong path, there's no explicit backtracking mechanism (Reflexion and ToT each filled this hole).
- One LLM call per step, making long tasks expensive (see Cost Control).
Influence on later systems
ReAct is the most-cited — and most often "reinvented" — Agent paradigm. LangChain's AgentExecutor, the OpenAI Agents SDK's runner, and Claude Code's main loop are all ReAct at heart. Shunyu Yao, one of the paper's authors, went on to participate in ToT, Reflexion, and SWE-agent — that lineage alone is a microcosm of Agent history. At the case level, compare SWE-agent and OpenHands.
4. Toolformer (2023)
Publication info: Schick et al., Meta AI. arXiv:2302.04761 (February 2023), published at NeurIPS 2023.
The problem it tackles
ReAct teaches the model to use tools via prompts, but prompts can only teach so much: the model doesn't truly "understand" when to call a tool, and switching toolsets means rewriting the prompt. Could the model teach itself to call external APIs?
Core method
Self-supervised learning, in three steps:
1. Sampling: given ordinary text, have the model insert candidate API calls
at positions where they "might help," e.g.:
"Pittsburgh is also known as [QA("What is Pittsburgh's nickname?")] the
Steel City"
2. Filtering: execute the API calls and compare the model's prediction loss
on the following text with and without the call result. Keep only calls
that actually reduce the loss (i.e., the result is genuinely useful).
3. Fine-tuning: fine-tune the model on the filtered corpus (the paper used
GPT-J 6.7B).The paper covered five tool types: calculator, QA system, search engine, translation system, and calendar.
Key experiments and data
The fine-tuned Toolformer significantly beat same-scale baselines across a range of downstream tasks, and learned to autonomously decide which tool to call, with what arguments, and how to weave results into its output — with no manually annotated call examples.
Limitations
- Tools can't be chained: one call can't depend on another's result; no composition.
- The sample-and-filter pipeline is expensive, and sensitive to "did the loss drop" as a proxy metric.
- The model is small (6.7B), a clear gap in the era of much larger models.
Influence on later systems
Toolformer proved "tool use can be trained into the weights." Months later OpenAI shipped function calling, then Gorilla, API-Bank, and native tool calling in every model family — all this road. Today the "prompt school" (ReAct lineage) and the "training school" (Toolformer lineage) have merged: foundation models ship with tool-calling built in, and Agent frameworks organize the calling loop.
5. Reflexion (2023)
Publication info: Shinn et al., Northeastern University & MIT & Princeton. arXiv:2303.11366 (March 2023), published at NeurIPS 2023.
The problem it tackles
Traditional RL lets Agents learn from trial and error, but sample efficiency is poor and it requires fine-tuning weights. The LLM-era question: can an Agent "learn from failure" without touching the weights?
Core method
Reflexion's answer: translate RL's reward into natural language, store it in memory, and read it back next round. The architecture has four roles:
┌─────────┐ generate trajectory ┌──────────┐
│ Actor │ ────────────────────→ │ Evaluator │ ← scores this round's performance
│ (LLM) │ │(rules/LLM)│
└─────────┘ └──────────┘
↑ │ feedback signal
│ ▼
┌─────────┐ ┌──────────┐
│ Episodic │ ←────────────────────│ Self- │ ← writes the cause of failure as
│ Memory │ stores reflection │Reflection│ a natural-language reflection
│(reflections)│ └──────────┘
└─────────┘
│
└── on the next attempt, the Actor reads the reflection text into its promptThe key insight: when humans fail, they don't "update gradients" — they jot down "last time I went wrong by not looking up the eastern sector first." An LLM happens to be an excellent consumer of exactly that kind of text.
Key experiments and data
- HumanEval code generation: pass@1 of 91%, beating GPT-4's single-shot 80% at the time — the paper's most famous number.
- Significantly outperformed no-reflection baselines on both ALFWorld (sequential decision-making) and HotpotQA (reasoning).
- Ablations showed that both the feedback signal's source (external environment vs. self-simulation) and form (scalar vs. free text) affect the gains.
Limitations
- Depends on a reliable evaluation signal: without unit tests or a verifiable environment, the Evaluator itself can misjudge, and reflection goes "in the wrong direction."
- Can fall into repeat-failure loops — similar reflections, unchanged behavior — so a max-round cap is needed as a backstop.
- The 91% on HumanEval was achieved with multiple attempts allowed, not strictly comparable to single-shot pass@1 — mind the framing when citing.
Influence on later systems
Reflexion made "self-correction" a standard Agent component: the "run tests → read errors → fix code" loop of coding agents (both SWE-agent and Devin), self-verification in RAG systems, and LLM-as-judge in evaluation systems are all intellectual descendants. It also inspired a wave of follow-up work in the "verbal RL" direction.
6. Generative Agents (2023)
Publication info: Park et al., Stanford University & Google Research. arXiv:2304.03442 (April 2023), published at ACM UIST 2023.
The problem it tackles
Earlier Agent research focused on "completing tasks"; this paper asked something entirely different: can an Agent live like a human? That is, produce believable human behavior — waking up, making breakfast, going to work, forming opinions, remembering yesterday, planning tomorrow.
Core method
The paper places 25 Agents in Smallville, a Sims-like sandbox town, each powered by a three-piece architecture:
┌────────────────────────────────────────────┐
│ Memory Stream │
│ Records every observation in chronological│
│ order, in natural language │
└────────────────────────────────────────────┘
│ retrieval ▲ write
▼ │
Retrieval score = recency
+ relevance (to the current situation, via vector similarity)
+ importance (scored by the LLM)
│
▼
┌─────────────┐ ┌─────────────┐
│ Reflection │ │ Planning │
│ (periodically│ │ (recursively │
│ distills │ │ generates a │
│ memories │ │ schedule, │
│ into higher-│ │ then breaks │
│ level insight)│ │ it into acts)│
└─────────────┘ └─────────────┘The most famous emergent case: give one Agent a single seed setting — "wants to host a Valentine's Day party" — and within two days, invitations spread through town on their own; Agents meet each other, arrange to attend together, and actually show up at the right place at the right time.
Key experiments and data
- Human evaluation found the full architecture's Agent behavior more "believable" than several ablated versions and human-written baselines.
- Ablations proved observation, planning, and reflection are all indispensable — remove reflection and the Agent can't remember who invited it; remove planning and the schedule falls apart.
Limitations
- Staggering cost: 25 Agents simulated for two days consumed considerable API budget even by the standards of the time; poor real-time performance.
- "Believability" is judged subjectively by human raters, with no objective metric; raters may even overestimate it because of the LLM's confident prose style.
- Agents make characteristic mistakes: misaligned memory retrieval, treating imaginings as fact, persona drift over long horizons.
Influence on later systems
This is the founding work of the "memory system" research direction. The memory stream + weighted retrieval (recency / relevance / importance) formula is still used directly by many memory frameworks (including LangMem and various RAG memory implementations), and "periodically reflecting to distill fragmentary memories into higher-level conclusions" is standard practice in today's long-term memory design. It also opened the branch of multi-agent social simulation (such as a16z's AI Town and various swarm simulations); see Multi-Agent Architecture.
7. Tree of Thoughts (2023)
Publication info: Yao et al., Princeton University & Google DeepMind. arXiv:2305.10601 (May 2023), published at NeurIPS 2023.
The problem it tackles
CoT is a straight line: generate left to right, no exploration, no backtracking. But many tasks (puzzles, planning, math proofs) are more like mazes — you need to try multiple paths, evaluate intermediate positions, and back up when stuck. Humans solve problems by searching, not by monologuing.
Core method
Model reasoning as search over a "tree of thoughts." At each step, generate multiple candidate thoughts; the LLM itself values each intermediate state (sure / maybe / impossible); then expand with a classic search algorithm:
┌─ T1 ── eval: maybe ─────┬─ T4 ── sure ──────→ keep expanding
Problem → step 1 ┼─ T2 ── eval: impossible ────→ ✂ prune
└─ T3 ── eval: sure ─────┬─ T5 ── impossible ─→ ✂
└─ T6 ── maybe ────→ backtrack, then expand
Search strategy: BFS (keep the best b states per step) or DFS (with backtracking)
Valuation: have the LLM score states independently, or vote over candidatesFour pluggable parts: how thoughts are decomposed, how thoughts are generated, how states are valued, and which search algorithm to use. This is what makes it more engineered than CoT — each part can be customized per task.
Key experiments and data
- Game of 24 (combine 4 numbers with arithmetic to make 24): GPT-4 + CoT solved only 4%; ToT reached 74%.
- Also substantially beat CoT on creative writing and 5×5 mini crosswords.
- The cost is equally striking: tens of times more LLM calls per problem than CoT — benefits and costs surge together.
Limitations
- Requires the task to be "decomposable with evaluable intermediate states"; open-ended tasks make thought granularity hard to define.
- The LLM grades its own states, so the evaluator is noisy; in situations demanding expert knowledge, the model often "doesn't know what it doesn't know."
- For most real business tasks, the 74%-vs-4% scenario is rare, and paying tens of times the cost for those cases usually isn't worth it.
Influence on later systems
ToT is the signature starting point of the "test-time compute" route. It was followed by Graph of Thoughts, LATS (combining MCTS with LLMs), and others; the deeper influence is conceptual — reasoning models like o1/R1 essentially internalized "search + evaluation" into the training process, letting the model learn on its own when to explore multiple paths. In engineering practice, ToT serves more as the theoretical reference for "how planning should search" than as a directly deployable solution.
8. Voyager (2023)
Publication info: Wang et al., NVIDIA & Caltech & UT Austin & Stanford & ASU (corresponding authors include Jim Fan and Anima Anandkumar). arXiv:2305.16291 (May 2023), formally published in TMLR (Transactions on Machine Learning Research), March 2024.
The problem it tackles
Earlier game Agents used RL to grind scores in fixed environments; change the task and you retrain. Voyager asked: can an LLM drive an agent that learns for life in an open world — exploring autonomously, accumulating skills, with capability compounding — all without human intervention and without fine-tuning weights?
Core method
Voyager runs in Minecraft (operated through the Mineflayer API), with three major components:
┌──────────────────┐
│ Automatic │ Proposes the next goal from the current state:
│ Curriculum │ "there's meat nearby → learn to hunt";
│ │ difficulty adapts and climbs with capability
└────────┬─────────┘
▼
┌──────────────────┐ On failure: read environment feedback/errors, fix the code
│ Iterative │ ─────────────────────────┐
│ Prompting │ │
│ │ ←──── self-verification: was the goal met?
└────────┬─────────┘
▼ On success: "package this executable code into the library"
┌──────────────────┐
│ Skill Library │ A skill = a documented JavaScript function,
│ │ reused via vector retrieval, composable into
│ │ more complex skills
└──────────────────┘Unlike Reflexion's "text reflections," Voyager's asset is executable code: once it learns "mine stone," that skill enters the library permanently, and the next task "craft a stone pickaxe" can call it directly. Skills are compositional, so capability snowballs.
Key experiments and data
- Compared with the previous SOTA: 3.3× more unique items obtained, 2.3× the travel distance, and up to 15.3× faster at unlocking key tech-tree milestones.
- Zero-shot generalization: carry the accumulated skill library into a brand-new Minecraft world and it solves new tasks from scratch, while other methods barely generalize.
- Throughout, it called GPT-4 as a black box without touching a single model parameter.
Limitations
- Strongly bound to environments like Minecraft where "state is programmatically readable and actions are programmatically executable"; the real world has no such clean API.
- The skill library accumulates broken skills (code that ran then, but breaks in new contexts), with no automatic cleaning mechanism.
- Cost grows linearly with exploration time — the bill for "lifelong learning" is also lifelong.
Influence on later systems
Voyager's "executable skill library" idea directly shaped code-agent design: "distill verified operations into reusable tools/scripts" in systems like OpenHands, and the skill/plugin mechanisms of various Agent frameworks, all carry its shadow. It is also the representative work of the "automatic curriculum learning + code as skill" paradigm — best read together with the CodeAct paper.
9. CodeAct (2024)
Publication info: Wang et al., UIUC. arXiv:2402.01030 (February 2024), published at ICML 2024.
The problem it tackles
After ReAct, agent actions were typically output as JSON or predefined text formats. This has two structural flaws: the action space is locked to a predefined toolset (tools can't compose, and slightly complex arguments get written wrong), and "calling an API" is far less natural for an LLM than "writing code" — the model's pretraining corpus is full of code, not JSON function calls.
Core method
Replace all predefined tools with one unified action space: let the Agent directly generate executable Python code.
Traditional JSON actions (one tool per step, composition is hard):
Action: {"tool": "search", "query": "Tesla stock price"}
Action: {"tool": "calculator", "expr": "..."}
CodeAct action (one block of code handles the composition logic):
```python
prices = [get_stock_price(t) for t in ["TSLA", "AAPL", "GOOG"]]
avg = sum(prices) / len(prices)
print(f"Average price: {avg:.2f}")
if avg > threshold:
send_alert(avg)
Code actions are executed by a Python interpreter, and the Agent, reading the output (errors included), can **dynamically revise its previous action** over multiple turns — control flow, variable reuse, exception handling, and library calls all come for free. The team also built CodeActInstruct, a 7k-example multi-turn CodeAct interaction fine-tuning dataset, and trained CodeActAgent on Llama 2 / Mistral.
### Key experiments and data
- Evaluating 17 LLMs on API-Bank and a newly built tool-use benchmark: CodeAct improved success rates by **up to about 20%** over JSON/text-format actions.
- Code actions have a hidden advantage: fewer interaction rounds to finish a task (one code block replaces a chain of JSON calls).
- Fine-tuned on a mix of CodeActInstruct and general instruction data, agent capability improved without harming general capabilities.
### Limitations
- Amplified security surface: executable code means sandbox isolation is mandatory, and injecting a malicious prompt can escalate to arbitrary code execution (see [Security](/advanced/security)).
- Heavy dependence on the interpreter environment: missing dependencies, timeouts, and side-effect cleanup are all engineering debt.
- On simple single-tool tasks, code actions actually cost more tokens.
### Influence on later systems
CodeAct essentially won the "action space wars." OpenHands' agents use bash/Python execution as their primary action, HuggingFace's smolagents made CodeAgent the default paradigm, and Computer Use products increasingly favor "give the model a terminal" over "give the model a pile of buttons." First author Xingyao Wang is a core author of OpenHands — the academic route grew directly into the product route.
## 10. SWE-agent (2024)
**Publication info**: Yang et al., Princeton University. arXiv:[2405.15793](https://arxiv.org/abs/2405.15793) (May 2024), published at NeurIPS 2024.
### The problem it tackles
By 2024, "hooking an LLM up to tools" was old news — yet the same model, with a different tool wrapper, could score several percentage points apart on SWE-bench. The paper's question: **is the interface itself a discipline?** Humans have IDEs; Agents, as "a new kind of terminal user," deserve interfaces designed for them.
### Core method
It proposes the **Agent-Computer Interface (ACI)** concept, with concrete design principles. SWE-agent's ACI includes:Design principle Concrete implementation ────────────────────────────────────────────────────────────── Actions should be "few but powerful" Dedicated commands instead of free-form bash: open / goto / scroll_down / edit (with line-number validation) Feedback should be "information-dense" After edit, immediately echo the changed context; reject execution on syntax errors and explain why Include guardrails humans don't get View only 100 lines at a time (prevents oversized output from flooding the context window) Docs built in Every failed command automatically attaches usage instructions
### Key experiments and data
- On SWE-bench (real GitHub issue fixing) it achieved **12.5% pass@1**, and **87.7%** on HumanEvalFix — both state of the art at the time (on GPT-4 as the base; the purely interactive Agent substantially beat non-interactive methods).
- Ablations showed every ACI design decision significantly affected scores — e.g., "no echo after editing" or "removing the line limit" both noticeably lowered success rates.
### Limitations
- 12.5% looks low today (the SWE-bench leaderboard was pushed up rapidly afterward); the paper's value is methodological, not numerical.
- The ACI design principles were distilled from trial and error, lacking a formal theory; the optimal ACI drifts with base-model capability.
- Evaluation is confined to software engineering; transferability to other domains needs separate validation.
### Influence on later systems
SWE-agent turned "designing tool interfaces for Agents" from a craft into a discipline: every coding Agent today ([Claude Code](/cases/claude-code), [Cursor](/cases/cursor), OpenHands) practices ACI thinking when designing toolsets — action granularity, echo policy, guardrail placement, everywhere are trade-offs discussed in the paper. It also evolved into an actively maintained open-source project; see the site's [SWE-agent case study](/cases/swe-agent).
## 11. Five Things These Papers Teach Us Together
**1. Reasoning should be externalized — but externalized doesn't mean correct.** CoT, ReAct, and ToT all share the premise "write the thinking out," yet all three hit the same wall: reasoning chains err, hallucinate, and self-deceive. So production systems never trust a single reasoning chain — ground it with tools (ReAct), search multiple paths (ToT), or verify after the fact (Reflexion). Saying only "CoT is powerful" in an interview won't pass.
**2. Model capability limits are compensated by system architecture.** MRKL made this clear in 2022: don't wait for an all-knowing model; build a system with a router + expert modules. Four years on, models are a hundred times stronger, and the judgment still holds — the "expert modules" have just been upgraded from calculators to MCP servers.
**3. Feedback loops are worth more than single-shot generation.** Reflexion's 91%, Voyager's snowballing skills, SWE-agent's test-driven fixes — all the gains come from the "execute → observe → correct" loop, not from cleverer one-shot prompts. When designing an Agent, first ask "where does the feedback signal come from," then "how should the prompt be written" — on tasks without reliable feedback, a reflection mechanism becomes an error-amplifying loop.
**4. Action-space design is a first-class engineering problem.** Toolformer proved tool use can be trained into weights, CodeAct proved code beats JSON, SWE-agent proved interface details decide success. Three generations of work converge on the same conclusion: don't toss off the action space like a config file — it matters as much as model selection.
**5. Memory is the watershed between an Agent as "tool" and an Agent as "system."** Generative Agents' memory stream, Reflexion's episodic memory, and Voyager's skill library correspond to three forms of memory: narrative memory, lesson memory, and capability memory. An Agent without memory reincarnates from scratch every time; once memory enters, retrieval, compression, forgetting, and contamination defense all become new problems — exactly what the [Memory](/components/memory) and [Context Engineering](/components/context-engineering) pages cover.
::: tip What to read next
With these 10 papers done, your "classics curriculum" in Agents is complete. For frontier directions (reasoning models, deep research, Computer Use, multi-agent systems), head to [Frontier Papers](/papers/frontier); for choosing papers by career goal, see [Reading Paths](/papers/paths); to put these ideas into code, start writing [Build an Agent Yourself](/practice/build-your-own).
:::
## References
- [Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (arXiv:2201.11903)](https://arxiv.org/abs/2201.11903) — the original CoT paper, NeurIPS 2022
- [MRKL Systems (arXiv:2205.00445)](https://arxiv.org/abs/2205.00445) — AI21 Labs' modular architecture manifesto
- [ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629)](https://arxiv.org/abs/2210.03629) — the foundational Agent Loop paper, ICLR 2023
- [Toolformer: Language Models Can Teach Themselves to Use Tools (arXiv:2302.04761)](https://arxiv.org/abs/2302.04761) — self-supervised tool learning, NeurIPS 2023
- [Reflexion: Language Agents with Verbal Reinforcement Learning (arXiv:2303.11366)](https://arxiv.org/abs/2303.11366) — verbal reinforcement learning, NeurIPS 2023
- [Generative Agents: Interactive Simulacra of Human Behavior (arXiv:2304.03442)](https://arxiv.org/abs/2304.03442) — memory stream architecture and the small-town experiment, UIST 2023
- [Tree of Thoughts: Deliberate Problem Solving with Large Language Models (arXiv:2305.10601)](https://arxiv.org/abs/2305.10601) — turning reasoning into tree search, NeurIPS 2023
- [SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (arXiv:2405.15793)](https://arxiv.org/abs/2405.15793) — the ACI design discipline, NeurIPS 2024