Skip to content

Reading Paths

At a glance A three-stage Agent paper reading list: 5 beginner papers to build intuition (ReAct, CoT, Toolformer, MRKL, Generative Agents), 8 intermediate papers to build system-level skills (Reflexion, Voyager, SWE-agent, CodeAct, and more), and a frontier stage focused on 2025-2026 agentic RL, Deep Research, and self-improvement. Each entry includes an arXiv ID, a one-line value statement, estimated reading time, and prerequisites.

Reading Paths ​

Papers in the Agent field have one defining trait: the volume is explosive, but only about twenty are genuinely worth reading word for word. Thousands of papers tagged "LLM agent" have appeared on arXiv since 2023, and most are variants, combinations, or evaluations of the same handful of core ideas. This reading list takes the opposite approach: first select the dozen-plus foundational papers for deep reading and build your conceptual framework; after that, you can classify any new paper with a single phrase — "it's a variant of X" — and a glance at its abstract is enough.

Three stages, each building on the last:

Stage 1 Beginner (5 papers)     Stage 2 Intermediate (8 papers)    Stage 3 Frontier (2025-2026)
Build intuition                 Build system skills                Track the newest directions
─────────────                   ─────────────────                  ──────────────────
CoT ──────────┐                 Reflexion / ToT                    Agentic RL (R1 and beyond)
ReAct ────────┼──►              Voyager / HuggingGPT ──►           Deep Research systems
MRKL          │                 CoALA / Survey                     Computer Use
Toolformer    │                 SWE-agent / CodeAct                Self-improvement (DGM)
Generative Ag.┘                                                    Multi-agent failure modes (MAST)
"What is an agent"              "How do you build an agent system" "Where is the field heading"

Reading advice

What matters isn't how many papers you read, but whether you can retell each one after finishing. The test: with the paper closed, can you explain in three sentences "what problem it solved, what the core mechanism is, and why it works"? If not, re-read the core sections before rushing to the next paper. For every paper, read it alongside the official code repo or the Paper Map.

1. How to Use This List ​

Each paper entry follows a uniform format with four pieces of information:

  • arXiv ID: all verified by search; just append the ID to arxiv.org/abs/ to access;
  • One-line value: the irreplaceable thing this paper contributed to the field;
  • Estimated reading time: based on "close-reading the core sections + skimming the experiments," excluding time spent running code;
  • Prerequisites: what to read first, if anything; omitted when there are none.

On the method of reading papers, a few well-worn but genuinely useful tips:

  1. Read the Introduction and Figure 1 first. The information density of an Agent paper is concentrated in its system architecture diagram; understand the figure and you're halfway there.
  2. In the experiments, focus on the ablations. The main results only tell you "it works"; the ablations tell you "why it works" — usually one component doing the heavy lifting, and that's the part you can take away.
  3. Watch out for benchmark contamination and survivorship bias. Many of the "massive gains" reported in 2023 were built on GPT-4 as the baseline and toy environments with a few dozen tasks. When you see a number, ask first: how big is the task set? Is the baseline fair?
  4. Skim at your own pace. This list has a clear dependency chain, but the order within a stage isn't strict — if you get stuck, skip ahead.

If you haven't yet read the site's What is an AI Agent and Anatomy of an Agent, read those two first — the terminology in the papers will flow much more easily.

2. Stage 1: Beginner — Building Intuition (5 papers) ​

These five papers answer one question: what is the minimal constitution of an agent. After reading them, you should be able to hand-roll a Thought-Action-Observation loop — which is also the theoretical wellspring of the Agent Loop section.

1. Chain-of-Thought Prompting (CoT) ​

  • arXiv: 2201.11903 (January 2022, Google, NeurIPS 2022)
  • One-line value: gets the model to write out its intermediate reasoning steps before answering — "think before you speak," a technique that seems obvious today, is the starting point of all agent reasoning capability.
  • Estimated reading time: 40 minutes (short main text; experiments can be skimmed)
  • Prerequisites: basic prompting skills

The point of reading this paper isn't "CoT improves scores" (long since common knowledge) but understanding why externalizing reasoning into text works: the reasoning trace in the context effectively serves as working memory, so the generation of subsequent tokens is conditioned on intermediate conclusions already written down. This "externalized thinking" intuition is the master key to understanding ReAct, Reflexion, and even reasoning models.

2. MRKL Systems ​

  • arXiv: 2205.00445 (May 2022, AI21 Labs)
  • One-line value: the earliest architecture to explicitly propose "LLM as router + external expert modules as executors" — the architectural prototype of every tool-use agent today.
  • Estimated reading time: 45 minutes

MRKL (Modular Reasoning, Knowledge and Language) is historically underrated. It predates ReAct by six months, and its core claim still holds: don't force an LLM to do what it's bad at (precise computation, real-time facts, private data) — route those to dedicated tools. Its weakness is that the engineering implementation (each expert had to be trained/fine-tuned) was made obsolete by the function calling era. Read it for the architectural idea; don't agonize over implementation details.

3. ReAct ​

  • arXiv: 2210.03629 (October 2022, Princeton + Google Brain, ICLR 2023)
  • One-line value: interleaves reasoning (Thought) and action (Action) in a single trace, letting the model "think, act, and observe as it goes" — the foundational work of the modern agent loop and the paper most likely to come up in interviews.
  • Estimated reading time: 1.5 hours
  • Prerequisites: CoT

If you read only one paper from this list, make it this one. ReAct's format is so plain it barely reads like a paper — just a loop of Thought: ... / Action: ... / Observation: ... — but it was the first to demonstrate that a reasoning trace can guide actions and that action feedback (observations) can in turn correct the reasoning, the two reinforcing each other. As you read, keep two questions in mind: how does it mitigate CoT's "making things up in a closed room" (hallucinated chains)? And how does it mitigate a pure action sequence's "undirected flailing"? Once these click, every later agent framework will feel like a variation on the same theme. Pairs well with the site's Planning section.

4. Toolformer ​

  • arXiv: 2302.04761 (February 2023, Meta AI, NeurIPS 2023)
  • One-line value: teaches a language model, via self-supervision, "when to call, which API to call, and with what arguments" — the first step from tool use as "prompt engineering" to tool use as "model capability."
  • Estimated reading time: 1 hour
  • Prerequisites: MRKL (to understand "why tools are needed")

Toolformer's approach: automatically insert API call candidates into text, keep only the calls that genuinely reduce the language model's loss, then fine-tune on that data. This idea of "using self-supervised signals to filter useful calls" foreshadowed later function calling fine-tuning and agentic RL. Its specific method (few-shot self-annotation) has been superseded by better approaches, but the judgment that "tool-use capability can be trained in, not just prompted out" shaped the whole industry's direction after 2024. See Tools and MCP.

5. Generative Agents ​

  • arXiv: 2304.03442 (April 2023, Stanford + Google, UIST 2023)
  • One-line value: 25 agents live autonomously for two days in a virtual small town and emergent social behavior appears — the design template for the memory-stream / reflection / planning trio, and essential reading on the origins of agent memory systems.
  • Estimated reading time: 1.5 hours
  • Prerequisites: none (though ReAct helps you appreciate its action layer)

This paper's value isn't "the town is fun" — it's that it delivers a complete cognitive architecture: observations enter a memory stream; retrieval is weighted by recency / relevance / importance; salient memories trigger reflection that generates higher-level abstractions; reflection then drives planning. Countless agent products (including many companion and game products) copied this design pattern outright. As you read, ask: which parts exist "for the demo effect," and which would "any long-running agent need"? The site's Memory section expands this machinery for engineering use.

Stage 1 self-check

After the five papers, try answering: what's the essential difference between ReAct's Thought and CoT's reasoning? Which route won — MRKL's router or ReAct's autonomous decision-making — and why? Which unaddressed problem in the ReAct framework does Generative Agents' memory stream solve? If you can answer all three fluently, Stage 1 is done.

3. Stage 2: Intermediate — Building System Skills (8 papers) ​

The beginner papers teach you "how to make a single agent run"; the intermediate papers answer a harder question: how to make an agent work reliably on long-horizon, realistic, failure-prone tasks. These eight cover reflection, search-based planning, lifelong learning, task orchestration, an architecture survey, and the two schools of software engineering agents.

6. Reflexion ​

  • arXiv: 2303.11366 (March 2023, Princeton et al., NeurIPS 2023)
  • One-line value: writes "lessons from failure" as natural-language reflections into memory and feeds them back into the prompt on the next trial — improving through experience purely in language, without touching model weights.
  • Estimated reading time: 1 hour
  • Prerequisites: ReAct

Reflexion calls itself "verbal reinforcement learning," but read with clear eyes: it isn't RL, there's no gradient update, and all the "learning" happens in context. That's both its strength (plug and play) and its ceiling (limited context length; reflection quality depends on the model's capacity for self-examination). Its Actor → Evaluator → Self-Reflection → Memory quartet became the template for nearly every later self-correction system. A question worth pondering alongside: when is reflection useful, and when is it just wasted tokens? (Hint: without a reliable feedback signal, reflection often reinforces hallucination.)

7. Tree of Thoughts (ToT) ​

  • arXiv: 2305.10601 (May 2023, Princeton + Google, NeurIPS 2023)
  • One-line value: extends single-chain reasoning into "generate multiple candidate thoughts → self-evaluate → tree search," turning planning from one-shot generation into a backtrackable explicit search.
  • Estimated reading time: 1 hour
  • Prerequisites: CoT

ToT's engineering value is often questioned (multiple samples plus per-level evaluation multiplies token costs several-fold), but its conceptual value is enormous: it was the first to transplant the classic AI view that "problem solving = search through a space of thoughts" fully into the LLM context, directly inspiring later reasoning models (the o1/R1 generation can be understood as "internalizing ToT's search into the model"). As you read, focus on how BFS/DFS couples with LLM self-evaluation, and why "evaluating a thought" and "generating a thought" need separate prompts.

8. HuggingGPT ​

  • arXiv: 2303.17580 (March 2023, Zhejiang University + Microsoft Research Asia, NeurIPS 2023)
  • One-line value: an LLM acts as controller, decomposing tasks and dispatching them to specialist models on HuggingFace — the flagship of the "LLM orchestrates expert models" route and an engineering template for multi-tool orchestration.
  • Estimated reading time: 1 hour
  • Prerequisites: MRKL

HuggingGPT's four-stage pipeline (task planning → model selection → task execution → response generation) is the best specimen for understanding "orchestration-type agents." It continues MRKL's lineage but swaps the "expert" from an API to an entire model community. Note its hard flaws as you read: one wrong planning step collapses the whole chain, and model description cards often mismatch real behavior — exactly why later agentic workflows emphasize error recovery and verification.

9. Voyager ​

  • arXiv: 2305.16291 (May 2023, NVIDIA et al., TMLR 2024)
  • One-line value: a lifelong-learning agent in Minecraft: an automatic curriculum proposes new goals, a skill library accumulates reusable code, and execution feedback drives iterative correction — the complete answer to "how does an agent keep getting stronger."
  • Estimated reading time: 1.5 hours
  • Prerequisites: ReAct, Generative Agents

Voyager is the dual of Generative Agents' ideas in the "capability growth" direction: one accumulates memories, the other skills. Its skill library (verified code snippets stored in a vector DB, retrieved and composed on demand) is the academic prototype of the "skill/tool accumulation" features in today's agent products — you can see its shadow in every vendor's Agent Skills mechanisms in 2025-2026. As you read, ask: why is self-verification before a skill enters the library critical? Without that gate, the library quickly gets polluted by broken code.

10. CoALA: Cognitive Architectures for Language Agents ​

  • arXiv: 2309.02427 (September 2023, Princeton)
  • One-line value: uses the cognitive-science framework of "cognitive architectures" to systematically map the agent design space: how many kinds of memory, how to divide the action space, how to organize the decision loop — not a paper proposing a new method, but a map that helps you classify every agent paper.
  • Estimated reading time: 2 hours
  • Prerequisites: all of Stage 1

CoALA rewards slow reading. It decomposes an agent's decision-making into a proposal → evaluation → selection → execution loop, splits memory into working / episodic / semantic / procedural, and then files hundreds of existing works into this framework. After reading it, your first reaction to a new paper will be "this is an improvement to module X in the CoALA framework" — exactly the "classification ability" we mean. Strongly recommend reading it side by side with the site's Paper Map.

11. A Survey on LLM-based Autonomous Agents ​

  • arXiv: 2308.11432 (August 2023, Renmin University of China (Gaoling School) et al., published in Frontiers of Computer Science)
  • One-line value: a classic survey from a Chinese team that breaks agent construction into four modules — Profile / Memory / Planning / Action — plus a review of evaluation methods — a reference book for filling gaps.
  • Estimated reading time: 2-3 hours (selective reading; no need to read it all)
  • Prerequisites: all of Stage 1

Use a survey like a dictionary, not a novel. Suggested approach: read the classification-framework chapters first, then check which box each of the nine papers you've read falls into; for boxes you've never heard of (multimodal perception, embodied directions, say), go read those sections. Its division of labor with CoALA: CoALA gives the cognitive perspective, this one the engineering-module perspective — you need both to have a complete conceptual coordinate system.

12. CodeAct ​

  • arXiv: 2402.01030 (February 2024, UIUC et al., ICML 2024)
  • One-line value: unifies the agent's action space from "one JSON function call at a time" into "executable Python code" — composable, with control flow and error handling — yielding significantly higher success rates.
  • Estimated reading time: 1 hour
  • Prerequisites: ReAct, Toolformer

CodeAct answers a very practical engineering question: what should the agent's "hands" look like? JSON function calling invokes one tool and returns one result per step; code natively supports loops, conditionals, intermediate variables, and try/except. This judgment deeply shaped later products — Claude Code, OpenHands, and various "Code Mode" offerings are all CodeAct-lineage at heart. As you read, also consider its cost: sandbox security for code execution (discussed in the site's Security section).

13. SWE-agent ​

  • arXiv: 2405.15793 (May 2024, Princeton, NeurIPS 2024)
  • One-line value: introduces the Agent-Computer Interface (ACI) concept — the commands, feedback formats, and editing tools an agent uses should be as carefully designed as a UI for humans — and used it to hit early SOTA on SWE-bench.
  • Estimated reading time: 1.5 hours
  • Prerequisites: ReAct; ideally knowing what SWE-bench is

SWE-agent's most important insight: beyond model capability, interface design is itself performance. With the same base model, designing "open file" as an open command with a numbered-line window, and editing as an edit command with syntax checks, changes success rates by a wide margin. This explains why every coding agent's toolset increasingly looks alike — everyone is converging on a good ACI. Anyone building a coding agent cannot skip this paper; read it with the site's SWE-agent case study to see the full path from paper to open-source project.

A common pitfall at this stage

Don't try to reproduce every paper's experiments. Stage 2 is about building a sense of the "design space": given a new agent system, you should be able to say whether its planning uses ToT or ReAct, what memory structure it uses, whether its action space is JSON or code, and whether its feedback loop is Reflexion-style or environment-driven. Save reproduction for the one or two papers you actually intend to build on (SWE-agent or Voyager are recommended — both have high-quality code).

4. Stage 3: Frontier — 2025-2026 Directions (Optional) ​

The first two stages are "settled consensus"; this stage is "ongoing debate." The papers/reports below were all verified by search, but frontier conclusions are fluid — when reading, focus on how problems are defined rather than on specific numbers.

14. DeepSeek-R1 and Agentic RL ​

  • arXiv: 2501.12948 (January 2025, DeepSeek)
  • One-line value: demonstrated that large-scale reinforcement learning (RLVR, reinforcement learning with verifiable rewards) can train long-chain reasoning into a model rather than prompting it out — internalizing much of what 2023 built with prompt tricks (CoT, part of ToT's motivation) into model capability, and opening the agentic RL mainline.
  • Estimated reading time: 1.5 hours
  • Prerequisites: CoT, ToT

The right way to read R1 is with a question in mind: once reasoning is trained into the model, what remains that needs to be externalized in an agent system? The answer: scaffolding for long-horizon tasks, tool orchestration, environmental feedback — which is precisely where agent engineering has centered since 2025. To follow this line systematically, continue with the survey The Landscape of Agentic Reinforcement Learning for LLMs (arXiv 2509.02547, September 2025), which maps the terrain of agentic RL task environments, reward design, and training methods.

15. Deep Research Agents: A Systematic Examination And Roadmap ​

  • arXiv: 2506.18096 (June 2025, survey)
  • One-line value: systematically dissects the tech stack of Deep Research products (the deep research features of OpenAI / Google / Perplexity): retrieval strategy, multi-step planning, long-context management, report generation and citation — the entry map for building research agents.
  • Estimated reading time: 1.5-2 hours
  • Prerequisites: ReAct, basic concepts of RAG

Deep Research was one of 2025's most successful new agent categories: the user poses a research question, the system autonomously retrieves dozens to hundreds of sources, and produces a long cited report. This survey cracks the product black box into researchable modules. As you read, compare against real products (OpenAI's deep research released February 2025, Google Gemini's same-named feature) and their actual performance, and think about which modules clearly still have room to improve — that's usually where papers are missing, which means opportunity.

16. OSWorld and Computer Use ​

  • arXiv: 2404.07972 (April 2024, HKU + Salesforce + CMU et al., NeurIPS 2024)
  • One-line value: the first benchmark evaluating agents in a real operating system environment: 369 cross-app tasks running in real VMs, where the agent can only see screenshots and send mouse/keyboard events — the measurement cornerstone of the computer-use direction.
  • Estimated reading time: 1 hour
  • Prerequisites: none (basic understanding of multimodal models suffices)

OSWorld is itself a benchmark paper, but the paradigm it defined (screenshot → action coordinates → execute → screenshot again) is the basic shape adopted by Anthropic's Claude computer use launched in October 2024 and OpenAI's Operator. The point of reading this paper is understanding why GUI agents are so hard: visual grounding, long-horizon state tracking, and the risk of irreversible operations. The success rates reported for early models were abysmal (single digits to low teens in percent); two years later the numbers have climbed a lot but are still far from "solved" — check the current leaderboard as you read to get a feel for how fast this direction is climbing.

17. Darwin-Gödel Machine (DGM) ​

  • arXiv: 2505.22954 (May 2025, Sakana AI + UBC + Vector Institute)
  • One-line value: lets a coding agent rewrite its own code, using empirical benchmark scores on SWE-bench and the like as evolutionary pressure to accumulate self-improvement in an open-ended way — the signature 2025 work in the "self-improving agent" direction.
  • Estimated reading time: 1.5 hours
  • Prerequisites: SWE-agent, Voyager

DGM is the de-theorized version of the Gödel Machine idea: it doesn't require "provably beneficial modifications" — anything "empirically effective on benchmarks" is kept in the population, with an open-ended evolution archive preventing local optima. Its results (self-improvement pushing SWE-bench scores from about 20% to about 50%; exact numbers per the paper version) demonstrate the engineering feasibility of recursive self-improvement while also exposing its risk surface — self-modifying agents are inherently harder to audit and constrain, so consider the security dimension as you read.

18. Why Do Multi-Agent LLM Systems Fail? (MAST) ​

  • arXiv: 2503.13657 (March 2025, UC Berkeley et al.)
  • One-line value: a systematic failure analysis of 7 mainstream multi-agent frameworks, distilling 14 failure modes (the MAST taxonomy) — a high-quality bucket of cold water on "multi-agent mania."
  • Estimated reading time: 1 hour
  • Prerequisites: basic understanding of multi-agent architecture

This paper's stance suits engineers perfectly: multi-agent systems' gains on benchmarks are often marginal, and most failures aren't the model's fault but system design problems — unclear task specifications, agent-to-agent conversations going off the rails, missing verification steps. After reading it, you'll have a far clearer judgment on "when to use multi-agent" (in most cases a single strong agent with good tools suffices; multi-agent is over-engineering). Its failure taxonomy also works directly as a debug checklist for your own system.

5. Choosing a Path by Goal ​

Not everyone needs all eighteen papers. Trim to your goal:

GoalMust readOptionalCan skip
Building a coding agent (job hunting or product)ReAct, CodeAct, SWE-agent, Reflexion, R1Toolformer, OSWorld, DGMGenerative Agents, HuggingGPT
Building a research/search agentReAct, CoT, Deep Research survey, ToTSurvey (2308.11432), MASTVoyager, OSWorld
Building a general agent platform/frameworkReAct, MRKL, Toolformer, CoALA, CodeActHuggingGPT, MAST, Agentic RL surveyGenerative Agents details
Building a multi-agent systemGenerative Agents, MAST, CoALAHuggingGPT, VoyagerSWE-agent, OSWorld
Interview prep (short on time)CoT, ReAct, Reflexion, SWE-agent, R1CoALA framework chaptersDefer everything else
Researching self-improvementReflexion, Voyager, DGMR1, Agentic RL surveyMRKL, HuggingGPT

A few supplementary suggestions:

  • Job-seeking readers: after the must-read column, go straight to the Job Knowledge Map and Interview Questions — organizing your answers around the concepts from these papers is better value than reading more papers.
  • Readers preparing to write papers / do research: Stage 3 is only the entrance; snowballing forward through each paper's citation network is the right path. Also see the site's Frontier Papers.
  • Everyone on a "building products" path: read Practice Pitfalls alongside the papers — papers won't tell you how to account for cost and latency.

6. What If You Just Can't Get Through Them ​

Some practical advice to close. Papers aren't comfortable reading for most engineers, so here are ways to lower the barrier:

  1. Watch a second-hand explainer before reading the original. Nearly all of these classic papers have close-reading videos on YouTube/Bilibili; watch 20 minutes to build the global picture, then read the original to fill in details — far more efficient.
  2. Run code instead of reading experiments. You can hand-write ReAct in 50 lines of Python (see Build an Agent Yourself); getting it running once beats reading the experiments section three times.
  3. Read with your own question in mind. Questions like "if I wanted to add memory to this agent, what would I change?" force you to read the architecture instead of the story.
  4. Give yourself permission to skip the math. The vast majority of these papers involve no heavy math; for the few that do (e.g., loss design in the RL parts), skipping on the first pass is perfectly fine.

References ​