Appearance
Reading Paths
Too many papers, too little time—that's the first hurdle in this field's literature. This page boils the thirty-plus papers in the Paper Map down to three routes: pick the one that matches the goal in front of you, read in the given order, and skip the curation. Each route comes with a rough time estimate and a pointer to where to go once you're done.
Quick Start (about 4 hours, build intuition)
- ReAct — understand the origin of the agent loop
- SWE-bench — understand what this field measures
- SWE-agent — understand how harness/ACI translates into scores
- One survey (pick any) — connect the dots into a full picture
All four have a topic home and a one-sentence contribution in the Paper Map; ReAct and SWE-agent additionally have paper-by-paper close readings.
Deep Dive (about 2 weeks, topic by topic)
Close-read 1–2 papers in each group, in this order: Reasoning and Acting → Tool Learning → Memory → Planning → Self-Improvement → Evaluation. After each group, come back to the matching core components page on this site to cross-check. Focus your comparison on: what tasks ReAct's alternating structure and ToT's search structure each suit; and which sits closer to production practice—MemGPT's paging approach or Generative Agents' reflection stream.
Putting It into Practice (about 1 week, read with questions in mind)
- Agentless — calibrate your "don't over-engineer" baseline first
- SWE-agent — learn interface design principles (how to constrain the action space, how to compress feedback)
- CodeAct — decide whether your action space uses JSON or code
- WebArena / OSWorld — learn how execution-based evaluation is put together, for self-testing
- Combine it with Build Your Own Harness and implement a minimal closed loop hands-on
A shared reading discipline
The absolute numbers in these papers (success rates, pass@1) depend almost entirely on the model versions and harness configurations of their time, so comparing them across papers is of limited value. What's worth taking away is the mechanism and failure analysis, not the scores. For more common questions, see Reading Discipline and FAQ.
Further Reading
- Paper Map — look up the coordinates of any paper by topic
- Classic Paper Close Readings — paper-by-paper close readings of 7 foundational papers
- Frontier — new work after 2024
- Reading Discipline and FAQ — how to read the numbers and what to mind when citing