Skip to content

Paper Map

At a glance Organizing the 2022-2026 Agent research landscape into a navigable map: nine sections covering reasoning and planning, tool use, memory, multi-agent, embodied AI, Coding Agents, evaluation, safety, and surveys — each with verified representative papers' arXiv IDs and a one-line status assessment.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Paper Map ​

From late 2022 to now, Agent research has gone from "readable in full" to "no one can read it all." The good news: no more than fifty works have real structural influence, and most of the rest are permutations and combinations of them. This map divides the landscape into nine sections, each listing only verified representative papers that are still cited repeatedly, along with a one-line status assessment of that direction.

Suggested usage: first read Reading Paths to pick an order, then use this map to locate sections, and finally go to Core Papers for paper-by-paper breakdowns. All arXiv IDs were verified by search (as of August 2026) and link straight to the abstract pages.

1. Reasoning & Planning ​

The foundation of all Agent capability. This line of work asks: how does the model "think" before it acts?

PaperarXivOne-line contribution
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models2201.11903A single "Let's think step by step"-style prompt unlocked multi-step reasoning
Self-Consistency Improves Chain of Thought Reasoning2203.11171Sample multiple reasoning chains and vote — the first mature scheme trading compute for accuracy
ReAct: Synergizing Reasoning and Acting in Language Models2210.03629Interleaves Thought/Action/Observation into a loop; the de facto starting point of the Agent paradigm
Reflexion: Language Agents with Verbal Reinforcement Learning2303.11366Generates written reflections after failure, stores them in memory, and retries the next trial carrying the lesson
Tree of Thoughts: Deliberate Problem Solving with LLMs2305.10601Extends single-chain reasoning into a searchable thought tree (BFS/DFS + self-evaluation)

Status in one line: prompt-level reasoning tricks (CoT/ToT) have been internalized into the weights of reasoning models like o1/R1 — you no longer need to hand-write chains of thought; but ReAct's "reasoning-action interleaving" skeleton and Reflexion's "RL in words" idea remain the default design of every Agent Loop. See Planning and Agent Loop.

Assessment

Reading ToT in 2026 requires a historical lens: what it proved is that "search + an evaluator" significantly improves reasoning, an idea inherited by later MCTS distillation and RL training (not prompt trees). In practice you should almost never run prompt-based ToT in a production Agent — it costs tens of times more than a single chain, and reasoning models have covered most of the benefit.

2. Tool Use ​

This line turns the LLM from a "chatterbox" into "a controller that can invoke external capabilities" — the watershed between Agent and chatbot.

PaperarXivOne-line contribution
TALM: Tool Augmented Language Models2205.12255Early proof that a small model + tool fine-tuning can overtake a pure large model
Toolformer: Language Models Can Teach Themselves to Use Tools2302.04761Self-supervised annotation of API call sites, letting the model teach itself "when to call, what to call"
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends2303.17580An LLM as controller dispatching expert models on Hugging Face — models as tools
Gorilla: Large Language Model Connected with Massive APIs2305.15334Retrieval-aware fine-tuning for massive APIs; the benchmark work on tool-call accuracy

Status in one line: tool calling is basically "solved by engineering" — mainstream models natively support function calling, and the MCP protocol unified the tool integration layer; the hard problems shifted from "can it call" to "permissions, idempotency, failure recovery." Toolformer-style training methods still matter, but mainly for model vendors rather than application engineers.

3. Memory & Personalization ​

Without memory, every conversation is an Agent's first day on the job. This line studies how information is written, organized, retrieved, and forgotten.

PaperarXivOne-line contribution
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (RAG)2005.11401The original framework of parametric memory + non-parametric retrieval — older than the Agent concept itself, yet the foundation of all external memory
Generative Agents: Interactive Simulacra of Human Behavior2304.03442Memory stream + importance/recency/relevance-scored retrieval + reflective abstraction; the template for Agent memory architectures
MemGPT: Towards LLMs as Operating Systems2310.08560Treats the context window as RAM and external storage as disk, letting the LLM manage its own paging in and out

Status in one line: RAG is infrastructure, not a research frontier; the real frontier is "memory as an operating system" — the MemGPT line evolved into products like Letta, while "letting the Agent autonomously decide what to remember and what to forget" still has no consensus best solution. For engineering trade-offs, see Memory Systems and RAG.

4. Multi-Agent ​

Once a single Agent's capability tops out, the natural idea is to form a team. This line splits into two branches: "simulating societies" and "collaborating on work."

PaperarXivOne-line contribution
CAMEL: Communicative Agents for "Mind" Exploration2303.17760Role-playing two-agent dialogue, using inception prompting to stay on track
Generative Agents (origin of the society-simulation branch)2304.0344225 Agents living autonomously in a small town; the breakout work on emergent behavior
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework2308.00352Encodes a software company's SOP into multi-Agent collaboration flows — dual constraints of roles + process
AgentVerse: Facilitating Multi-Agent Collaboration2308.10848A composable multi-Agent framework with four stages: recruit-discuss-execute-evaluate
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation2308.08155Abstracts everything into "conversable Agents + conversation patterns"; Microsoft's multi-Agent programming framework

Status in one line: the society-simulation branch has academic volume but little industrial adoption; the collaborate-on-work branch split into frameworks (AutoGen, CrewAI) and products (Manus, etc.). Community consensus is also returning to sanity — multi-Agent is over-engineering in most scenarios, and the capability boundary of a single Agent + good tools extends much further than imagined in 2023. For a deeper discussion, see Multi-Agent Architecture.

5. Embodied & Game ​

Games and robots are the best proving grounds for Agents: closed environments, measurable rewards, cheap failures.

PaperarXivOne-line contribution
Do As I Can, Not As I Say (SayCan)2204.01691The LLM supplies "what to do," the affordance function supplies "what can be done"; multiply them to pick the action
Voyager: An Open-Ended Embodied Agent with Large Language Models2305.16291A lifelong-learning Agent in Minecraft: automatic curriculum + skill library + self-verification
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control2307.15818Represents robot actions as text tokens and trains VLA models together with internet knowledge

Status in one line: Voyager's "skill library = executable code memory" directly inspired later Coding Agents; SayCan's affordance idea evolved into VLA foundation models (RT-2 and successors). Embodied intelligence became a funding and publication hotspot in 2025-2026, but its engineering stack has clearly diverged from software Agents.

6. Coding Agent ​

The first section of Agent research to close the commercial loop. Code has an executable validation signal (did the tests pass?), making it a natural fit for the Agent Loop.

PaperarXivOne-line contribution
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?2310.06770Used real GitHub issues to define the standard exam for "automated bug fixing"
Executable Code Actions Elicit Better LLM Agents (CodeAct)2402.01030Lets the Agent output executable code as its action — stronger than JSON tool calls
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering2405.15793Introduces the ACI concept: a command-line interface designed for Agents matters as much as one designed for humans
OpenHands: An Open Platform for AI Software Developers as Generalist Agents2407.16741The open-source general software Agent platform; the engineering culmination of CodeAct's ideas

Status in one line: this is the fastest "paper → product" conversion among the nine sections — SWE-bench (and its hand-curated subset SWE-bench Verified) became the arms-race leaderboard for every coding product, and the ideas of SWE-agent and OpenHands flowed directly into Claude Code, Cursor, and Devin. Case breakdowns: SWE-agent, OpenHands, Claude Code.

Caution

Watch for a reporting trap when reading Coding Agent papers: the SWE-bench scores papers report usually include retries and scaffold tuning, and can differ severalfold from what you get running the same model bare. When comparing two systems, first confirm they use the same benchmark variant and the same retry budget.

7. Evaluation & Benchmarks ​

Agent evaluation is the most "honest" section of this field: it keeps proving that problems we thought were solved actually aren't.

PaperarXivOne-line contribution
AgentBench: Evaluating LLMs as Agents2308.03688The first systematic Agent evaluation suite across 8 types of interactive environments (OS, databases, web, games, etc.)
WebArena: A Realistic Web Environment for Building Autonomous Agents2307.13854Tests web manipulation on self-hosted copies of real websites, dragging web Agent evaluation from toys to live fire
GAIA: a Benchmark for General AI Assistants2311.12983Conceptually simple but tool-chain-dependent multi-step assistant tasks: humans score 92%, the GPT-4-plugins of the time under 20%
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments2404.07972Execution inside a real OS, graded by system state; the standard exam hall for computer-use Agents
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains2406.12045Uses an LLM-simulated user in multi-turn adversarial testing, focused on rule adherence and consistency; famously shows success rate (pass^k) collapsing under retries

Status in one line: the core contradiction of benchmarking has shifted from "no exam questions" to "questions being gamed" — vendors' leaderboard-targeted optimization caused score inflation, and the community's answer is a steady stream of harder, contamination-resistant variants (SWE-bench Verified, τ-bench follow-ups, etc.). For methodology on building your own evals, see Evaluation Systems and Evals in Practice.

8. Safety & Alignment ​

Once an Agent has hands and feet, the risk of "saying the wrong thing" escalates into "doing the wrong thing." This is the youngest and most underrated of the nine sections.

PaperarXivOne-line contribution
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents2406.13352A dynamic arena for prompt injection attack/defense: measures both "task completion rate" and "compromise rate after injection"
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents2410.09024110 multi-step malicious tasks measure misuse risk; found that leading models are "surprisingly compliant" with malicious agentic requests, and simple jailbreak templates transfer to Agent settings

Status in one line: the research consensus is harsh — wrap a safety-aligned chatbot in an Agent shell and safety does not automatically carry over; tool-returned content is the largest injection surface. No silver bullet on defense; engineering relies on the "boring but effective" combo of least privilege, human confirmation, and trace auditing. For a practical checklist, see Security and Defense.

Boundary

This section has far fewer papers than the other eight — not because the problem is small, but because disclosing attack/defense details is risky, and much of the work is published as vendor system cards rather than papers. Beyond the literature, be sure to follow vendors' security documentation and red-team reports.

9. Surveys ​

For a quick global view, surveys are the shortcut — but their shelf life is usually about a year, so read the new ones, not the old.

Paper / ArticleSourceOne-line contribution
A Survey on Large Language Model based Autonomous Agents2308.11432Proposed the Profile / Memory / Planning / Action four-module split, widely adopted by later surveys
The Rise and Potential of Large Language Model Based Agents: A Survey2309.07864Organizes the landscape through the lens of "single Agent → multi-Agent → Agent society"; enormous — use as a dictionary
LLM Powered Autonomous Agents (Lilian Weng blog)lilianweng.github.ioNot a paper, but its Planning + Memory + Tool use trichotomy is the most-cited mental model in the community

Status in one line: surveys after 2024 turned toward specialties (memory surveys, multi-Agent surveys, evaluation surveys, reasoning-model surveys), and the value of catch-all mega-surveys is declining — a sign of a maturing field. Beginners should still start with Lilian Weng's blog post; half an hour to build the skeleton.

10. Knowledge Graph: Relationships and Evolution Across Themes ​

The diagram below connects the nine sections into a web along time and influence. Arrows show "the flow of ideas": sources on the left, downstream on the right.

                        2020        2022           2023                    2024           2025-26
                          │           │              │                       │               │
  ┌─────────────┐   RAG ──┴──► CoT ───┴──► ReAct ────┼──► Reflexion ────────┼──► reasoning ─┴──► internalized
  │ ① Reasoning │        (2005.11401) (2201.11903) (2210.03629)           │    models       into model weights
  │  & Planning │                          │          (ToT 2305.10601 branch: prompt tree search,   (o1/R1 line)
  └─────────────┘                          │           later displaced by RL training)
                                           ▼
  ┌─────────────┐   TALM ──────► Toolformer ──► HuggingGPT / Gorilla ──► function calling ──► MCP
  │ ② Tool Use  │  (2205.12255)  (2302.04761)   (2303.17580 / 2305.15334)   native support    protocol
  └─────────────┘                          │                                                  unification
                                           ▼
  ┌─────────────┐   Generative Agents ◄─────┘   MemGPT ──────────────► Letta and other memory products
  │ ③ Memory    │   (2304.03442, memory stream)  (2310.08560, LLM-as-OS)
  └─────────────┘          │
                           ▼
  ┌─────────────┐   CAMEL ──► MetaGPT / AgentVerse / AutoGen ──► multi-Agent frameworks ──► back to
  │ ④ Multi-    │ (2303.17760) (2308.00352 / 2308.10848 / 2308.08155)   boom            single Agent
  │   Agent     │                                                                       (a rational rethink)
  └─────────────┘
  ┌─────────────┐   SayCan ──► Voyager ──► RT-2 (VLA) ──────────► embodied AI as its own track
  │ ⑤ Embodied  │ (2204.01691) (2305.16291) (2307.15818)
  │   & Game    │                │
  └─────────────┘                │ skill library = code memory
                                 ▼
  ┌─────────────┐   SWE-bench ──► CodeAct ──► SWE-agent ──► OpenHands ──► Claude Code / Devin
  │ ⑥ Coding    │  (2310.06770)  (2402.01030) (2405.15793)  (2407.16741)   and other products
  └─────────────┘        │
                         ▼
  ┌─────────────┐   AgentBench ──► WebArena / GAIA ──► OSWorld / τ-bench ──► contamination-
  │ ⑦ Evaluation│  (2308.03688)  (2307.13854 / 2311.12983) (2404.07972 / 2406.12045) resistant variants
  └─────────────┘                                                        │
  ┌─────────────┐   AgentDojo / AgentHarm ───────────────────────────────┘ attack/defense evals
  │ ⑧ Safety    │  (2406.13352 / 2410.09024)                                               become routine
  └─────────────┘
  ┌─────────────┐   2308.11432 / 2309.07864 / Lilian Weng blog ──► era of specialized surveys
  │ ⑨ Surveys   │
  └─────────────┘

How to read this diagram:

  1. Horizontal is time: before 2022, RAG stood alone; 2023 was the Cambrian explosion — seven of the nine sections took shape that year; 2024 shifted the focus to evaluation and Coding engineering; 2025-2026 saw reasoning models absorbing section ①'s prompt tricks, plus an arms race in safety and evaluation.
  2. Vertical is dependency: without ①'s ReAct there is no ②'s tool-calling loop; without ③'s memory template there is no ④'s society simulation; ⑤'s skill-library idea directly nourished ⑥. Coding Agent is the biggest "sink" in the whole graph — almost every section's results get cashed in there.
  3. Vanished branches matter too: prompt-based ToT and multi-Agent social simulation were both loud directions that were engineering-debunked or shelved. Papers absent from this map mostly died along those roads.

If you're job hunting or changing careers, sections ⑥ and ⑦ come up in interviews most often — they test both depth (SWE-bench's scoring mechanism) and breadth (the applicability boundaries of different benchmarks). Companion reading: High-Frequency Interview Questions, Frontier.

References ​