Skip to content

Start Here

At a glance Entry point to the paper close-reading section: pick a reading path that matches your goal, or jump straight into the paper map, classic close readings, and the frontier.

Paper Close Readings: Start Here ​

Every component of an agent harness—the loop, tools, memory, planning, evaluation—has a prototype that was validated by experiments in a specific paper. This section organizes the relevant research into four ways to use it: read by path, close-read paper by paper, look up by topic, and follow the frontier.

Where Should You Start? ​

  1. Only have two hours → Take the "Quick Start" route in Reading Paths: ReAct → SWE-bench → SWE-agent → one survey, to build domain intuition.
  2. Want to put it into practice → Take the "Putting It into Practice" route in Reading Paths, with an emphasis on the close readings of Agentless, SWE-agent, and CodeAct, then head back to Build Your Own Harness and start building.
  3. Want to do research or write a survey → Start by flipping through the Paper Map to get your bearings, then work through the topic groups, and finally skim Frontier to see which direction the open problems have moved since 2025.
  4. Just want to look up one paper → Locate it by topic in the Paper Map; every entry lists an arXiv ID, the institution, the year, and a one-sentence contribution.

What Else Is in This Section? ​

  • Reading Paths — three reading routes ordered by goal (4 hours / 2 weeks / 1 week)
  • Classic Paper Close Readings — seven papers, one close reading each: ReAct, MRKL, Toolformer, Reflexion, Generative Agents, Voyager, and SWE-agent—the problem, the mechanism, the experimental findings, and what they imply for harness design
  • Paper Map — the full academic map grouped by topic, with a timeline of paradigm shifts
  • Frontier — six research threads from 2025–2026: long-horizon execution, context engineering, memory infrastructure, multi-agent systems, self-improvement, and evaluation discipline
  • Reading Discipline and FAQ — how to read the numbers in papers, and what to watch out for when citing them

A word of advice

Whichever route you take, hold on to one thing first: the absolute scores in these papers depend almost entirely on the model versions and harness configurations of their time. What's worth taking away is the mechanism and the failure analysis, not the leaderboard numbers. For details, see the FAQ.