Skip to content

Reading Paths

At a glance Reading orders for three kinds of readers: a 4-hour quick start, a 2-week deep dive, and a 1-week route into engineering practice. Each path notes what to read, why, and where to head once you finish.

Reading Paths ​

Too many papers, too little time—that's the first hurdle in this field's literature. This page boils the thirty-plus papers in the Paper Map down to three routes: pick the one that matches the goal in front of you, read in the given order, and skip the curation. Each route comes with a rough time estimate and a pointer to where to go once you're done.

Quick Start (about 4 hours, build intuition) ​

  1. ReAct — understand the origin of the agent loop
  2. SWE-bench — understand what this field measures
  3. SWE-agent — understand how harness/ACI translates into scores
  4. One survey (pick any) — connect the dots into a full picture

All four have a topic home and a one-sentence contribution in the Paper Map; ReAct and SWE-agent additionally have paper-by-paper close readings.

Deep Dive (about 2 weeks, topic by topic) ​

Close-read 1–2 papers in each group, in this order: Reasoning and Acting → Tool Learning → Memory → Planning → Self-Improvement → Evaluation. After each group, come back to the matching core components page on this site to cross-check. Focus your comparison on: what tasks ReAct's alternating structure and ToT's search structure each suit; and which sits closer to production practice—MemGPT's paging approach or Generative Agents' reflection stream.

Putting It into Practice (about 1 week, read with questions in mind) ​

  1. Agentless — calibrate your "don't over-engineer" baseline first
  2. SWE-agent — learn interface design principles (how to constrain the action space, how to compress feedback)
  3. CodeAct — decide whether your action space uses JSON or code
  4. WebArena / OSWorld — learn how execution-based evaluation is put together, for self-testing
  5. Combine it with Build Your Own Harness and implement a minimal closed loop hands-on

A shared reading discipline

The absolute numbers in these papers (success rates, pass@1) depend almost entirely on the model versions and harness configurations of their time, so comparing them across papers is of limited value. What's worth taking away is the mechanism and failure analysis, not the scores. For more common questions, see Reading Discipline and FAQ.

Further Reading ​