Skip to content

The Paper Map

At a glance An academic map of Agent Harness research: representative papers grouped by topic, plus a timeline of paradigm shifts, to help you place any new paper on the grid.

The Paper Map ​

Agent Harness — an agent's skeletal frame — wasn't the invention of any single paper. It is a systemic picture that dozens of papers assembled piece by piece over three years: ReAct supplied the interleaved reasoning-and-acting main loop, Toolformer proved tool use could be trained into the model, SWE-bench dragged evaluation into real software engineering, and SWE-agent put interface design in a paper title for the first time. This page organizes that work into a map by topic, so you can figure out where a new paper lands on the grid and what to read first.

How to use this map

If you only have two hours, go straight to Reading Paths and pick one. Every paper is tagged with its arXiv ID, and the publication years and author institutions have been verified against the original publication pages. Within each topic, papers are ordered chronologically so you can watch the ideas evolve.

Paradigm Shift Timeline ​

text
2021          2022                2023                              2024
 │             │                   │                                 │
 HumanEval     CoT ─────────────►  ReAct sets the Agent Loop paradigm │
 (Codex)       SayCan             Toolformer / Gorilla / ToolLLM      │
               Inner Monologue    Generative Agents / MemGPT          │
               STaR               Voyager / Reflexion / Self-Refine   │
                                  WebArena / AgentBench / GAIA ──────►│
                                  SWE-bench ─────────────────────────►│
                                                                    OSWorld
                                                                    SWE-agent (ACI)
                                                                    CodeAct
                                                                    AutoCodeRover / Agentless
                                                                    OpenHands

Three clear phase shifts:

  1. 2022: from generating answers to interleaved decisions. CoT got models to reason before answering; ReAct wove reasoning and external actions together, and the Agent Loop took its first recognizable shape.
  2. 2023: the harness components get invented one by one. Tools (Toolformer, Gorilla), memory (Generative Agents, MemGPT), skill libraries and self-reflection (Voyager, Reflexion), and real-environment evaluation (WebArena, SWE-bench) all surfaced within roughly the same year — by its end, the anatomy of the harness was essentially drawn.
  3. 2024: system design and evaluation discipline take center stage. SWE-agent introduced the Agent-Computer Interface (ACI) concept, Agentless cross-examined the need for agents from the outside with an agent-free pipeline, and OSWorld pushed evaluation into real operating systems. The research focus shifted from making models think better to fencing them in better — which is exactly the heart of harness research.

Reasoning and Acting Paradigms ​

This group of papers defines the agent's heartbeat: what the model thinks at each step, what it does, and how the two alternate.

PaperInstitutionYearContribution in one sentence
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsGoogle Research2022Uses few-shot exemplars to elicit intermediate reasoning steps, establishing the think-before-you-answer baseline capability
ReAct: Synergizing Reasoning and Acting in Language ModelsPrinceton / Google Brain2022Interleaves reasoning traces with external actions — the template for nearly every modern agent loop
Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsPrinceton / Google DeepMind2023Expands single-chain reasoning into a searchable, backtrackable tree of thoughts, lifting the Game of 24 success rate from 4% to 74%
Executable Code Actions Elicit Better LLM Agents (CodeAct)UIUC2024Replaces JSON tool calls with executable Python code as a unified action space, improving success rates by up to 20%

Independent judgment

ReAct's real legacy isn't the prompt template — it's that the observe-think-act turn structure was cemented as the harness's main-loop protocol. CodeAct reminds us that the representation of the action space (JSON versus code) is itself a harness design decision, and a consequential one. Both papers deserve a word-by-word read.

Tool Learning ​

How does a model learn to call tools — through prompt orchestration, or by internalizing them through training?

PaperInstitutionYearContribution in one sentence
Toolformer: Language Models Can Teach Themselves to Use ToolsMeta AI2023Self-supervised: the model learns when to call an API and what arguments to pass, with only a handful of examples per API
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceZhejiang University / Microsoft Research Asia2023An LLM acts as the controller dispatching expert models on Hugging Face, with language as the universal interface
Gorilla: Large Language Model Connected with Massive APIsUC Berkeley2023Retrieval-augmented fine-tuning beats GPT-4 on API call generation while easing API hallucination
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsTsinghua University / ModelBest2023The ToolBench dataset plus a depth-first-search decision tree give open-source models command of over ten thousand real APIs

Tool learning cuts both ways for the harness: the better the harness writes its tool descriptions, the more accurately the model calls them (see Tools & MCP); and the better the model gets trained at calling tools, the less orchestration the harness has to do.

Self-Improvement ​

How does an agent get better at tasks without gradient updates?

PaperInstitutionYearContribution in one sentence
STaR: Bootstrapping Reasoning With ReasoningStanford2022Fine-tunes the model on its own rationales that lead to correct answers, bootstrapping reasoning ability
Reflexion: Language Agents with Verbal Reinforcement LearningNortheastern / MIT / Princeton2023Replaces weight updates with verbal reflection, reaching 91% HumanEval pass@1 (as reported in the paper)
Self-Refine: Iterative Refinement with Self-FeedbackCMU / Allen Institute for AI2023A single LLM doubles as generator, feedback provider, and refiner, iterating on outputs without any training
Voyager: An Open-Ended Embodied Agent with Large Language ModelsNVIDIA / Caltech / UT Austin2023Automatic curriculum plus an executable code skill library plus self-verification in Minecraft, delivering lifelong learning

A note on reading Reflexion

The 91% HumanEval figure Reflexion reports was obtained under conditions that allow iterative attempts and test feedback — it is not on the same track as the 80% single-shot baseline. That gap is precisely the harness's value, but cite the numbers with their conditions spelled out.

Memory ​

What an agent remembers, how it retrieves it, and how it compresses it determine how long a task it can run.

PaperInstitutionYearContribution in one sentence
Generative Agents: Interactive Simulacra of Human BehaviorStanford / Google Research2023An observe-reflect-plan memory stream architecture from which 25 simulated townsfolk emerged with believable social behavior
MemGPT: Towards LLMs as Operating SystemsUC Berkeley2023Borrows tiered storage and "virtual context management" from operating systems, using interrupts to page memory in and out on its own

Voyager's skill library can also be read as a kind of memory — except that what it stores isn't facts but things the agent knows how to do. All three routes (retrieval-based memory streams, OS-style paging, programmatic skill libraries) have counterparts in today's production harnesses; see Memory Systems.

Planning ​

Separating "what to do next" from the model itself is the most persistent debate in agent research.

PaperInstitutionYearContribution in one sentence
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan)Google Research2022Constrains LLM planning with affordance scores over skills, so plans land on things the robot can actually do
Inner Monologue: Embodied Reasoning through Planning with Language ModelsGoogle Research2022Closed-loop language feedback (success detection, scene description, human input) acts as an inner monologue to improve embodied planning
On the Planning Abilities of Large Language ModelsArizona State University2023A critical evaluation built on International Planning Competition domains: LLMs autonomously produce executable plans only about 3% of the time
LLM+P: Empowering Large Language Models with Optimal Planning ProficiencyUT Austin2023The LLM translates the natural-language problem into PDDL, hands it to a classical planner for the optimal solution, and translates the result back

Taken together, these papers tell one complete story: LLMs plan poorly on their own (Valmeekam et al.), but the picture improves dramatically once you outsource planning to a symbolic system (LLM+P) or constrain it in a closed loop with environment feedback (SayCan, Inner Monologue). This is the academic source of the harness's division-of-labor idea, mirrored in Planning & Task Decomposition.

Evaluation Benchmarks ​

Without benchmarks there is no harness engineering. The benchmarks below define the yardstick for whether an agent works.

BenchmarkInstitutionYearContribution in one sentence
HumanEval (Evaluating Large Language Models Trained on Code)OpenAI2021164 hand-written programming problems plus execution-based judging — the de facto starting point of code-generation evaluation
WebArena: A Realistic Web Environment for Building Autonomous AgentsCMU2023A self-hosted environment of fully functional real websites; the strongest GPT-4 agent of its day managed only 14.41% (humans: 78.24%)
AgentBench: Evaluating LLMs as AgentsTsinghua University2023Multi-dimensional evaluation across 8 interactive environments, systematically charting the agent-capability gap between commercial and open-source models
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Princeton20232,294 real GitHub issues, pushing evaluation from writing functions to fixing real repositories — the industry's gold standard
GAIA: a benchmark for General AI AssistantsMeta AI / Hugging Face et al.2023466 real-world questions that are easy for humans and hard for AI, calling for browsing, reasoning, and tool combinations
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer EnvironmentsUniversity of Hong Kong / CMU / Salesforce et al.2024369 open-ended tasks in real operating systems with execution-based evaluation; the best model scores only 12.24%

How to read a benchmark paper

Don't stop at the leaderboard numbers. Look at three things: where the tasks come from (realism), how success is judged (execution-based or match-based), and what the failure cases look like (which decides what the harness should supply). The "Claude 2 solved only 1.96%" figure in the SWE-bench paper matters because it proved the problem back then wasn't model intelligence — it was the mode of interaction.

Software Engineering Agents ​

Software engineering is the field where harness thinking has landed most completely, and it is the center of gravity of this site.

PaperInstitutionYearContribution in one sentence
SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringPrinceton2024Introduces the Agent-Computer Interface (ACI) concept, showing that interface design alone significantly moves agent performance
AutoCodeRover: Autonomous Program ImprovementNational University of Singapore2024Organizes code search around program structure (AST/classes/methods) rather than a flat file view, at a markedly lower cost than its contemporaries
Agentless: Demystifying LLM-based Software Engineering AgentsUIUC2024A three-stage localize-repair-verify agent-free pipeline that reached the top ranks of SWE-bench Lite at a fraction of the cost
OpenHands: An Open Platform for AI Software Developers as Generalist AgentsUIUC / CMU et al.2024A community-driven open platform: code execution sandboxes, multi-agent coordination, and integrated evaluation across 15 benchmarks
Why Agentless is this site's most important cautionary tale

Agentless deliberately forgoes the agent loop — the model never freely decides the next step, a fixed pipeline does all the work — and still beat every open-source software agent of its time. The lesson is not that "agents are useless"; it is that in harness design, every degree of freedom must have a reason. Autonomy piled on blindly gets paid for in reliability, cost, and debuggability. This dovetails with Design Principles and Common Pitfalls on this site.

Surveys ​

To build the full picture quickly, these two surveys complement each other.

PaperInstitutionYearContribution in one sentence
A Survey on Large Language Model based Autonomous AgentsRenmin University of China2023Proposes a unified construction-application-evaluation framework for organizing agent research, with a continuously maintained paper repository
The Rise and Potential of Large Language Model Based Agents: A SurveyFudan University2023A brain-perception-action three-part framework, plus agent societies and their philosophical roots; a long read

Further Reading ​

References ​