Appearance
Paper Map
From late 2022 to now, Agent research has gone from "readable in full" to "no one can read it all." The good news: no more than fifty works have real structural influence, and most of the rest are permutations and combinations of them. This map divides the landscape into nine sections, each listing only verified representative papers that are still cited repeatedly, along with a one-line status assessment of that direction.
Suggested usage: first read Reading Paths to pick an order, then use this map to locate sections, and finally go to Core Papers for paper-by-paper breakdowns. All arXiv IDs were verified by search (as of August 2026) and link straight to the abstract pages.
1. Reasoning & Planning
The foundation of all Agent capability. This line of work asks: how does the model "think" before it acts?
| Paper | arXiv | One-line contribution |
|---|---|---|
| Chain-of-Thought Prompting Elicits Reasoning in Large Language Models | 2201.11903 | A single "Let's think step by step"-style prompt unlocked multi-step reasoning |
| Self-Consistency Improves Chain of Thought Reasoning | 2203.11171 | Sample multiple reasoning chains and vote — the first mature scheme trading compute for accuracy |
| ReAct: Synergizing Reasoning and Acting in Language Models | 2210.03629 | Interleaves Thought/Action/Observation into a loop; the de facto starting point of the Agent paradigm |
| Reflexion: Language Agents with Verbal Reinforcement Learning | 2303.11366 | Generates written reflections after failure, stores them in memory, and retries the next trial carrying the lesson |
| Tree of Thoughts: Deliberate Problem Solving with LLMs | 2305.10601 | Extends single-chain reasoning into a searchable thought tree (BFS/DFS + self-evaluation) |
Status in one line: prompt-level reasoning tricks (CoT/ToT) have been internalized into the weights of reasoning models like o1/R1 — you no longer need to hand-write chains of thought; but ReAct's "reasoning-action interleaving" skeleton and Reflexion's "RL in words" idea remain the default design of every Agent Loop. See Planning and Agent Loop.
Assessment
Reading ToT in 2026 requires a historical lens: what it proved is that "search + an evaluator" significantly improves reasoning, an idea inherited by later MCTS distillation and RL training (not prompt trees). In practice you should almost never run prompt-based ToT in a production Agent — it costs tens of times more than a single chain, and reasoning models have covered most of the benefit.
2. Tool Use
This line turns the LLM from a "chatterbox" into "a controller that can invoke external capabilities" — the watershed between Agent and chatbot.
| Paper | arXiv | One-line contribution |
|---|---|---|
| TALM: Tool Augmented Language Models | 2205.12255 | Early proof that a small model + tool fine-tuning can overtake a pure large model |
| Toolformer: Language Models Can Teach Themselves to Use Tools | 2302.04761 | Self-supervised annotation of API call sites, letting the model teach itself "when to call, what to call" |
| HuggingGPT: Solving AI Tasks with ChatGPT and its Friends | 2303.17580 | An LLM as controller dispatching expert models on Hugging Face — models as tools |
| Gorilla: Large Language Model Connected with Massive APIs | 2305.15334 | Retrieval-aware fine-tuning for massive APIs; the benchmark work on tool-call accuracy |
Status in one line: tool calling is basically "solved by engineering" — mainstream models natively support function calling, and the MCP protocol unified the tool integration layer; the hard problems shifted from "can it call" to "permissions, idempotency, failure recovery." Toolformer-style training methods still matter, but mainly for model vendors rather than application engineers.
3. Memory & Personalization
Without memory, every conversation is an Agent's first day on the job. This line studies how information is written, organized, retrieved, and forgotten.
| Paper | arXiv | One-line contribution |
|---|---|---|
| Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG) | 2005.11401 | The original framework of parametric memory + non-parametric retrieval — older than the Agent concept itself, yet the foundation of all external memory |
| Generative Agents: Interactive Simulacra of Human Behavior | 2304.03442 | Memory stream + importance/recency/relevance-scored retrieval + reflective abstraction; the template for Agent memory architectures |
| MemGPT: Towards LLMs as Operating Systems | 2310.08560 | Treats the context window as RAM and external storage as disk, letting the LLM manage its own paging in and out |
Status in one line: RAG is infrastructure, not a research frontier; the real frontier is "memory as an operating system" — the MemGPT line evolved into products like Letta, while "letting the Agent autonomously decide what to remember and what to forget" still has no consensus best solution. For engineering trade-offs, see Memory Systems and RAG.
4. Multi-Agent
Once a single Agent's capability tops out, the natural idea is to form a team. This line splits into two branches: "simulating societies" and "collaborating on work."
| Paper | arXiv | One-line contribution |
|---|---|---|
| CAMEL: Communicative Agents for "Mind" Exploration | 2303.17760 | Role-playing two-agent dialogue, using inception prompting to stay on track |
| Generative Agents (origin of the society-simulation branch) | 2304.03442 | 25 Agents living autonomously in a small town; the breakout work on emergent behavior |
| MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework | 2308.00352 | Encodes a software company's SOP into multi-Agent collaboration flows — dual constraints of roles + process |
| AgentVerse: Facilitating Multi-Agent Collaboration | 2308.10848 | A composable multi-Agent framework with four stages: recruit-discuss-execute-evaluate |
| AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation | 2308.08155 | Abstracts everything into "conversable Agents + conversation patterns"; Microsoft's multi-Agent programming framework |
Status in one line: the society-simulation branch has academic volume but little industrial adoption; the collaborate-on-work branch split into frameworks (AutoGen, CrewAI) and products (Manus, etc.). Community consensus is also returning to sanity — multi-Agent is over-engineering in most scenarios, and the capability boundary of a single Agent + good tools extends much further than imagined in 2023. For a deeper discussion, see Multi-Agent Architecture.
5. Embodied & Game
Games and robots are the best proving grounds for Agents: closed environments, measurable rewards, cheap failures.
| Paper | arXiv | One-line contribution |
|---|---|---|
| Do As I Can, Not As I Say (SayCan) | 2204.01691 | The LLM supplies "what to do," the affordance function supplies "what can be done"; multiply them to pick the action |
| Voyager: An Open-Ended Embodied Agent with Large Language Models | 2305.16291 | A lifelong-learning Agent in Minecraft: automatic curriculum + skill library + self-verification |
| RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | 2307.15818 | Represents robot actions as text tokens and trains VLA models together with internet knowledge |
Status in one line: Voyager's "skill library = executable code memory" directly inspired later Coding Agents; SayCan's affordance idea evolved into VLA foundation models (RT-2 and successors). Embodied intelligence became a funding and publication hotspot in 2025-2026, but its engineering stack has clearly diverged from software Agents.
6. Coding Agent
The first section of Agent research to close the commercial loop. Code has an executable validation signal (did the tests pass?), making it a natural fit for the Agent Loop.
| Paper | arXiv | One-line contribution |
|---|---|---|
| SWE-bench: Can Language Models Resolve Real-World GitHub Issues? | 2310.06770 | Used real GitHub issues to define the standard exam for "automated bug fixing" |
| Executable Code Actions Elicit Better LLM Agents (CodeAct) | 2402.01030 | Lets the Agent output executable code as its action — stronger than JSON tool calls |
| SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering | 2405.15793 | Introduces the ACI concept: a command-line interface designed for Agents matters as much as one designed for humans |
| OpenHands: An Open Platform for AI Software Developers as Generalist Agents | 2407.16741 | The open-source general software Agent platform; the engineering culmination of CodeAct's ideas |
Status in one line: this is the fastest "paper → product" conversion among the nine sections — SWE-bench (and its hand-curated subset SWE-bench Verified) became the arms-race leaderboard for every coding product, and the ideas of SWE-agent and OpenHands flowed directly into Claude Code, Cursor, and Devin. Case breakdowns: SWE-agent, OpenHands, Claude Code.
Caution
Watch for a reporting trap when reading Coding Agent papers: the SWE-bench scores papers report usually include retries and scaffold tuning, and can differ severalfold from what you get running the same model bare. When comparing two systems, first confirm they use the same benchmark variant and the same retry budget.
7. Evaluation & Benchmarks
Agent evaluation is the most "honest" section of this field: it keeps proving that problems we thought were solved actually aren't.
| Paper | arXiv | One-line contribution |
|---|---|---|
| AgentBench: Evaluating LLMs as Agents | 2308.03688 | The first systematic Agent evaluation suite across 8 types of interactive environments (OS, databases, web, games, etc.) |
| WebArena: A Realistic Web Environment for Building Autonomous Agents | 2307.13854 | Tests web manipulation on self-hosted copies of real websites, dragging web Agent evaluation from toys to live fire |
| GAIA: a Benchmark for General AI Assistants | 2311.12983 | Conceptually simple but tool-chain-dependent multi-step assistant tasks: humans score 92%, the GPT-4-plugins of the time under 20% |
| OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments | 2404.07972 | Execution inside a real OS, graded by system state; the standard exam hall for computer-use Agents |
| τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains | 2406.12045 | Uses an LLM-simulated user in multi-turn adversarial testing, focused on rule adherence and consistency; famously shows success rate (pass^k) collapsing under retries |
Status in one line: the core contradiction of benchmarking has shifted from "no exam questions" to "questions being gamed" — vendors' leaderboard-targeted optimization caused score inflation, and the community's answer is a steady stream of harder, contamination-resistant variants (SWE-bench Verified, τ-bench follow-ups, etc.). For methodology on building your own evals, see Evaluation Systems and Evals in Practice.
8. Safety & Alignment
Once an Agent has hands and feet, the risk of "saying the wrong thing" escalates into "doing the wrong thing." This is the youngest and most underrated of the nine sections.
| Paper | arXiv | One-line contribution |
|---|---|---|
| AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents | 2406.13352 | A dynamic arena for prompt injection attack/defense: measures both "task completion rate" and "compromise rate after injection" |
| AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents | 2410.09024 | 110 multi-step malicious tasks measure misuse risk; found that leading models are "surprisingly compliant" with malicious agentic requests, and simple jailbreak templates transfer to Agent settings |
Status in one line: the research consensus is harsh — wrap a safety-aligned chatbot in an Agent shell and safety does not automatically carry over; tool-returned content is the largest injection surface. No silver bullet on defense; engineering relies on the "boring but effective" combo of least privilege, human confirmation, and trace auditing. For a practical checklist, see Security and Defense.
Boundary
This section has far fewer papers than the other eight — not because the problem is small, but because disclosing attack/defense details is risky, and much of the work is published as vendor system cards rather than papers. Beyond the literature, be sure to follow vendors' security documentation and red-team reports.
9. Surveys
For a quick global view, surveys are the shortcut — but their shelf life is usually about a year, so read the new ones, not the old.
| Paper / Article | Source | One-line contribution |
|---|---|---|
| A Survey on Large Language Model based Autonomous Agents | 2308.11432 | Proposed the Profile / Memory / Planning / Action four-module split, widely adopted by later surveys |
| The Rise and Potential of Large Language Model Based Agents: A Survey | 2309.07864 | Organizes the landscape through the lens of "single Agent → multi-Agent → Agent society"; enormous — use as a dictionary |
| LLM Powered Autonomous Agents (Lilian Weng blog) | lilianweng.github.io | Not a paper, but its Planning + Memory + Tool use trichotomy is the most-cited mental model in the community |
Status in one line: surveys after 2024 turned toward specialties (memory surveys, multi-Agent surveys, evaluation surveys, reasoning-model surveys), and the value of catch-all mega-surveys is declining — a sign of a maturing field. Beginners should still start with Lilian Weng's blog post; half an hour to build the skeleton.
10. Knowledge Graph: Relationships and Evolution Across Themes
The diagram below connects the nine sections into a web along time and influence. Arrows show "the flow of ideas": sources on the left, downstream on the right.
2020 2022 2023 2024 2025-26
│ │ │ │ │
┌─────────────┐ RAG ──┴──► CoT ───┴──► ReAct ────┼──► Reflexion ────────┼──► reasoning ─┴──► internalized
│ ① Reasoning │ (2005.11401) (2201.11903) (2210.03629) │ models into model weights
│ & Planning │ │ (ToT 2305.10601 branch: prompt tree search, (o1/R1 line)
└─────────────┘ │ later displaced by RL training)
▼
┌─────────────┐ TALM ──────► Toolformer ──► HuggingGPT / Gorilla ──► function calling ──► MCP
│ ② Tool Use │ (2205.12255) (2302.04761) (2303.17580 / 2305.15334) native support protocol
└─────────────┘ │ unification
▼
┌─────────────┐ Generative Agents ◄─────┘ MemGPT ──────────────► Letta and other memory products
│ ③ Memory │ (2304.03442, memory stream) (2310.08560, LLM-as-OS)
└─────────────┘ │
▼
┌─────────────┐ CAMEL ──► MetaGPT / AgentVerse / AutoGen ──► multi-Agent frameworks ──► back to
│ ④ Multi- │ (2303.17760) (2308.00352 / 2308.10848 / 2308.08155) boom single Agent
│ Agent │ (a rational rethink)
└─────────────┘
┌─────────────┐ SayCan ──► Voyager ──► RT-2 (VLA) ──────────► embodied AI as its own track
│ ⑤ Embodied │ (2204.01691) (2305.16291) (2307.15818)
│ & Game │ │
└─────────────┘ │ skill library = code memory
▼
┌─────────────┐ SWE-bench ──► CodeAct ──► SWE-agent ──► OpenHands ──► Claude Code / Devin
│ ⑥ Coding │ (2310.06770) (2402.01030) (2405.15793) (2407.16741) and other products
└─────────────┘ │
▼
┌─────────────┐ AgentBench ──► WebArena / GAIA ──► OSWorld / τ-bench ──► contamination-
│ ⑦ Evaluation│ (2308.03688) (2307.13854 / 2311.12983) (2404.07972 / 2406.12045) resistant variants
└─────────────┘ │
┌─────────────┐ AgentDojo / AgentHarm ───────────────────────────────┘ attack/defense evals
│ ⑧ Safety │ (2406.13352 / 2410.09024) become routine
└─────────────┘
┌─────────────┐ 2308.11432 / 2309.07864 / Lilian Weng blog ──► era of specialized surveys
│ ⑨ Surveys │
└─────────────┘How to read this diagram:
- Horizontal is time: before 2022, RAG stood alone; 2023 was the Cambrian explosion — seven of the nine sections took shape that year; 2024 shifted the focus to evaluation and Coding engineering; 2025-2026 saw reasoning models absorbing section ①'s prompt tricks, plus an arms race in safety and evaluation.
- Vertical is dependency: without ①'s ReAct there is no ②'s tool-calling loop; without ③'s memory template there is no ④'s society simulation; ⑤'s skill-library idea directly nourished ⑥. Coding Agent is the biggest "sink" in the whole graph — almost every section's results get cashed in there.
- Vanished branches matter too: prompt-based ToT and multi-Agent social simulation were both loud directions that were engineering-debunked or shelved. Papers absent from this map mostly died along those roads.
If you're job hunting or changing careers, sections ⑥ and ⑦ come up in interviews most often — they test both depth (SWE-bench's scoring mechanism) and breadth (the applicability boundaries of different benchmarks). Companion reading: High-Frequency Interview Questions, Frontier.
References
- ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629) — the paradigm starting point of the Agent Loop, the citation anchor of the whole map.
- Generative Agents: Interactive Simulacra of Human Behavior (arXiv:2304.03442) — the shared origin of the memory and multi-agent sections.
- Voyager: An Open-Ended Embodied Agent with LLMs (arXiv:2305.16291) — the representative work of the game section; its skill-library idea influenced Coding Agents.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (arXiv:2310.06770) — the starting gun of the Coding Agent arms race.
- GAIA: a Benchmark for General AI Assistants (arXiv:2311.12983) — the standard exam for general assistant capability.
- τ-bench: A Benchmark for Tool-Agent-User Interaction (arXiv:2406.12045) — multi-turn tool interaction and rule-adherence evaluation; its pass^k finding has had lasting influence.
- AgentDojo (arXiv:2406.13352) — the standard environment for prompt injection attack/defense evaluation.
- LLM Powered Autonomous Agents — Lilian Weng — the most-cited mental model for Agents (blog post).