Appearance
Frontier Developments
Core papers tell the 2022–2024 story: the components of a harness were invented one by one—the loop, tools, memory, planning, interfaces. After 2025, the research question flipped: the components all exist, but can they work together reliably on tasks that run hours or days? Failure analysis has replaced capability demos, context has been promoted from a prompting trick to a systems discipline, and the harness itself has even become an object of optimization and comparison.
This page lays out seven threads by theme. Each gets representative work (the problem, a one-line take on the method, and what it means for harness design) plus this site's independent judgment. All arXiv IDs and publication dates have been verified.
text
2025.02 A-MEM ──────────────────────── Memory gets structured: notes + links + evolution
2025.03 METR long-task measurement ─── "how long it can last" replaces "how often it's right" as the yardstick
2025.03 MAST ───────────────────────── a taxonomy of multi-agent failure modes (the first bucket of cold water)
2025.04 SICA ───────────────────────── agents start editing their own code
2025.05 Lost in multi-turn ─────────── average 39% score drop in multi-turn dialogue
2025.05 DGM ────────────────────────── open-ended evolution: SWE-bench 20% → 50%
2025.06 Anthropic multi-agent system ─ industry owns up to multi-agent's costs and limits
2025.07 Context Rot / Manus ────────── context engineering goes mainstream
2025.09 Illusion of diminishing returns ─ long-horizon failures live in execution, not reasoning
2025.10 ACE ────────────────────────── context as playbook, self-evolving
2025.12 RLM / Confucius CCA ────────── context as environment; the scaffold's value gets quantified
2026.05 Harness disclosure manifesto ─ scores without harness disclosure aren't comparable
2026.08 Agent Lightning / LEGO-RL ──── training moves into the skeleton (Harness-Native RL)Long-Horizon Agents: From "Getting It Right" to "Going the Distance"
Taken together, these works move the yardstick of agent capability from "per-task success rate" to "time horizon", and explain why long tasks are hard.
METR's long-task measurement (Measuring AI Ability to Complete Long Tasks, arXiv:2503.14499, March 2025). Problem: how do you actually measure a model's agentic capability? Method: ignore scores and measure the "50%-task-completion time horizon"—the length of human tasks a model can complete autonomously with 50% reliability. They found this horizon has been doubling roughly every 7 months since 2019, and frontier models in early 2025 sat around the 1-hour mark. Why it matters for harnesses: this is a harness-sensitive metric—under different context-management and error-recovery designs, the same model's "how long it can hold up" differs completely, which makes time horizon a more honest touchstone for a harness than single-step accuracy.
The Illusion of Diminishing Returns (The Illusion of Diminishing Returns, arXiv:2509.09677, September 2025). Problem: do slowing gains on short-task benchmarks mean scaling has hit its ceiling? Method: separate "executing a task" from "reasoning about a task"—hand the model the knowledge and plan it needs, and measure only its ability to carry out long instruction sequences. Three findings matter enormously for harness design:
- Tiny gains in single-step accuracy compound into exponential growth in the length of tasks a model can complete—the "diminishing returns" on short benchmarks are a measurement illusion;
- Long-horizon failures come mostly from execution errors, not from a shortage of reasoning ability;
- There is a self-conditioning effect: when the context contains the model's own errors from earlier turns, the odds of future mistakes rise significantly. Scaling the model alone doesn't remove the effect, but thinking (long-chain reasoning) mitigates it substantially.
text
The compounding effect of per-step reliability (why 1% is a big deal)
95% per-step success rate × 100 consecutive steps ≈ 0.6% overall success rate
99.9% per-step success rate × 1000 consecutive steps ≈ 36.8% overall success rate
→ On long tasks, every error the harness intercepts and every trajectory
it cleans up does work on that exponential curveThe direct implication of self-conditioning for context engineering
If "the model's own errors sitting in the context" are toxic in themselves, then a harness that compresses, rewrites, or summarizes trajectories should not faithfully preserve the full history of failed attempts—the trajectory cleanup in context engineering (collapsing redundant error output, replacing raw error dumps with structured summaries) now has theoretical backing from measurement research.
Vending-Bench (arXiv:2502.15840, February 2025) sketches the failure shape of long horizons from the other side: the model runs a virtual vending machine over a long stretch, and what's tested is long-term coherence. Even the strongest models fall into "delusional spirals" after long runs—clinging to a false belief (say, that a shipment never arrived) and building an ever more absurd chain of decisions around it. Anthropic later had Claude run an actual vending machine (Project Vend, June 2025) and observed the same kind of long-run instability. The implication for harnesses is blunt: long tasks need externally anchored verification and state write-back. Counting on the model to "think its way clear" a few hundred turns in is not reliable.
Read together, the three papers prescribe the same cure for harness design: stopping and budget controls in the agent loop, continuous cleanup of "toxic" error trajectories out of context, critical state externalized where the model can't tamper with it, and risky operations on long tasks routed through the permission gate. None of these is solved by "waiting for a stronger model".
Context Engineering Becomes a Discipline
The most important conceptual shift of 2025 was "context engineering" going from an in-joke among a few practitioners to industry vocabulary. Two long-form industry posts defined it: the Manus team's Context Engineering for AI Agents: Lessons from Building Manus (July 2025) and Anthropic's Effective Context Engineering for AI Agents (September 2025). The former contributed many of the now widely cited techniques: keeping the KV-cache prefix stable, using repeated rewrites of todo.md to "recite" the goal into near-term attention (recitation), and treating the filesystem as an infinitely large external context; the latter systematized the principle of "minimal sufficient context". Details in context engineering.
Academia supplied two key pieces of evidence in the same window.
Context Rot (Chroma research report, July 2025, research.trychroma.com/context-rot). Problem: as context windows grow, does performance actually follow? Method: controlled measurement across 18 mainstream models, varying input length while holding task difficulty fixed. Finding: model performance degrades monotonically as input tokens increase, even when the task gets no harder—there is a systematic gap between the advertised "1M-token window" and "effective context". Why it matters for harnesses: stuffing more into the context is not free; curation is a responsibility the harness cannot duck.
LLMs Get Lost in Multi-Turn Conversation (arXiv:2505.06120, Microsoft Research / Salesforce, May 2025). Problem: do models perform as well in multi-turn conversation as in a single turn? Method: 200,000+ simulated conversations comparing single-turn full-information prompts with multi-turn sharded disclosure. Finding: every model tested dropped 39% on average in the multi-turn setting, and the decomposition shows the damage is mostly a reliability collapse (variance explodes) rather than a capability drop—the model makes hasty assumptions in early turns, locks in an answer too early, then refuses to backtrack. Why it matters for harnesses: long-session agents must explicitly persist intermediate conclusions (files, checklists, plan artifacts), because the model's "implicit memory" across turns is systematically unreliable—that is the measurement-grade evidence behind the "externalize state" argument in planning.
Recursive Language Models (arXiv:2512.24601, MIT CSAIL, December 2025) offer a radical answer. Problem: very long inputs are doomed to crush the context window. Method: don't feed the long prompt to the transformer—park it in a REPL environment as a variable, and let the model write code to inspect it, slice it, and fire recursive subcalls at itself. The paper reports that RLMs handle inputs two orders of magnitude beyond the window, and on four long-context tasks they clearly beat compression baselines and common coding scaffolds at comparable cost. Why it matters for harnesses: this pushes the subagent idea to its extreme—context management is environment design, and a long input stops being "content to digest" and becomes "an environment to explore".
Memory: From Memory Streams to Memory Infrastructure
The 2023 memory papers (Generative Agents and MemGPT) invented the mechanisms; the 2025 memory papers turned those mechanisms into subsystems that are governed, retrievable, and evolvable. Three papers, three routes:
| System | Core mechanism in one line | How memory is organized | Implication for harness design |
|---|---|---|---|
| A-MEM (arXiv:2502.12110) | Borrowing from the Zettelkasten card box: each memory becomes a structured note (keywords, tags, contextual description), semantic links are built automatically, and old memories evolve as new experience arrives | A dynamically evolving network of notes | Memory entries need relationships among themselves, so retrieval can carry context along with it |
| Mem0 (arXiv:2504.19413) | Extracts, updates, and deletes candidate facts from conversations; a production-grade memory layer built mainly on vector retrieval | A flat fact store + graph extension (Mem0ᵍ) | The paper reports it beats the full-context approach on the LOCOMO benchmark at lower latency and token cost—memory isn't a feature, it's a cost structure |
| MemOS (arXiv:2505.22101) | Treats memory as a resource an operating system manages: MemCube uniformly wraps parametric, activation, and plaintext memory, with metadata and governance attributes (lifecycle, permissions, audit) | Memory cubes + a scheduling framework | Memory needs governance: expiry, permissions, and provenance are all harness responsibilities |
One shared thread
The three papers disagree on how to organize memory but agree on something else: the write path needs criteria—what is worth storing, and when to update or delete. This continues the lineage of Toolformer's "keep it only if useful" filter: a memory store without write criteria misleads more the more you retrieve from it. Implementation details in memory systems.
Multi-Agent: The Hype and the Cold Water
Multi-agent research in 2025 has a rare quality: its two most influential publications are, respectively, an industry "field report" and an academic "study of failure".
Anthropic's multi-agent research system (How we built our multi-agent research system, June 2025). This is the engineering retrospective of the Claude Research feature: an orchestrator-worker structure where the lead agent decomposes the task, spawns subagents to search in parallel, then aggregates. The most valuable part of the post is its honest boundary statement: by its own numbers, a single agent burns roughly 4x the tokens of an ordinary chat, and a multi-agent system roughly 15x—multi-agent only pays off when the task is high-value, parallelizable, and the subtasks' information needs don't overlap.
MAST: a taxonomy of multi-agent failure modes (Why Do Multi-Agent LLM Systems Fail?, arXiv:2503.13657, March 2025). Problem: why do multi-agent systems actually fail? Method: failure attribution across 7 mainstream multi-agent frameworks and 200+ annotated trajectories, distilled into 14 failure modes in 3 categories (specification and system-design flaws, inter-agent misalignment, and task verification and termination issues). One finding runs through everything: verification is the weak link—results get passed along unchecked, and errors amplify down the chain.
Read the two side by side and this site's judgment is: for now, multi-agent architecture is better understood as an expensive way to parallelize than as a multiplier of intelligence. What it buys is throughput (parallel search, parallel exploration); what it doesn't buy is reliability (coordination and verification costs go up, not down). The engineering default remains "single agent + a good harness + constrained delegation of subtasks"—which is exactly the route Claude Code's subagent design takes: subagents exist to isolate context and explore in parallel, not to simulate a "team".
Self-Improvement: The Harness Becomes the Optimization Target
Reflexion in 2023 let an agent improve its behavior through reflection; the self-improving agents of 2025 improve the agent's own code and context—the harness went from "a fixed structure written by humans" to "a variable under optimization".
SICA (A Self-Improving Coding Agent, arXiv:2504.15228, April 2025). Problem: can an agent improve itself without a human tweaking its prompts and tools? Method: drop the "meta-agent improves target agent" hierarchy and let the agent edit its own codebase directly, with the edit history serving as a memory of improvements. It was one of the earliest complete closed-loop demonstrations.
Darwin Gödel Machine (arXiv:2505.22954, May 2025, Sakana AI et al.). Problem: can self-improvement be open-ended evolution rather than local patching? Method: maintain an archive of agents; each round, sample one version from the archive, have the foundation model generate interesting variants, validate them empirically on a coding benchmark, and admit the survivors—a Darwinian open-ended search. The results turn heads: SWE-bench climbed from 20.0% to 50.0%, Polyglot from 14.2% to 30.7%, and the improvements evolution found were precisely harness-level ones—better code-editing tools, long-context window management, peer-review mechanisms. Worth noting: DGM ran inside a sandbox under human supervision the whole way.
ACE (Agentic Context Engineering, arXiv:2510.04618, Stanford / SambaNova et al., October 2025) charts a route that never touches code. Problem: without changing weights or architecture, can a system improve itself by evolving the context itself? Method: treat the system prompt as a continuously evolving playbook, with three roles—generator, reflector, curator—adding and pruning entries over time, deliberately countering brevity bias and context collapse. The paper reports average gains of about 10.6% and 8.6% over strong baselines on AppWorld agent tasks and financial-analysis tasks—no weights touched, all of it context iteration.
The verification loop for self-improvement must live outside
The shared premise of these three works matches the lesson of Reflexion: whether an improvement is real is settled by external benchmark / environment signals, not by the model's own feelings. DGM gates archive admission on SWE-bench; ACE distills entries from task feedback. "Self-improvement" without a reliable verifier is just adding fuel to hallucination—the recurring pattern in common pitfalls. For a panoramic view, see the survey of self-evolving agents (arXiv:2507.21046, July 2025).
The inference for harness engineering is far-reaching: a harness's code, prompts, and tool descriptions are all objects that can be automatically searched and optimized—the prior advantage of hand-written harnesses is fading, while the machinery for "telling whether an improvement is real" (reliable local benchmarks, trajectory audits) becomes the new moat. See observability.
Training Moves into the Skeleton: Harness-Native RL
In mid-August 2026 a dense cluster of papers landed on arXiv that turn the harness from "a shell around inference" into "an environment for training"—the model is no longer merely wrapped in a harness; it gets trained inside one. At least three papers in three days:
Agent Lightning v1.0 (Towards Harnessed Agentic RL, arXiv:2608.17528, August 18, 2026). Problem: running RL on an agent usually means rewriting the agent's code around a training framework, and every harness swap means starting over. Method: a decoupled architecture—an LLM endpoint proxy plugs any agent into the RL training loop without touching the agent code. The paper states its first-generation scheme has been adopted by training frameworks including verl Uni-Agent, AReaL 2.0, slime, and Polar.
LEGO-RL (Harness-Native Reinforcement Learning for Coding Agents, arXiv:2608.17393, same day). Problem: RL for coding agents increasingly depends on real harnesses running for long stretches, but the harness's native execution environment and policy-gradient training are misaligned by nature—environment crashes and reward hacking pollute the outcome signal, and train/inference mismatch distorts rollouts. Method: build the training harness-native from the ground up, aligning training with inference inside the real harness.
ClawGym II (arXiv:2608.16798, August 17) takes a third path: no harness changes, no peeking at internal state—treat the whole skeleton as a black-box environment for RL, attacking the problem that training "through a complex harness is hard to scale" on long-horizon tasks.
Independent take
May's harness-disclosure manifesto proved that "the same model scores differently under different harnesses"; August's cluster is the logical next step: if the harness determines performance, let the model learn inside the most faithful harness there is. The implication for practitioners is direct—the harness you write is no longer just a deployment asset; it starts to shape what the model can learn to become, and harness stability (crash rate, signal cleanliness) turns from an engineering metric into a training metric.
Three more from the same week are worth a skim: Demystifying Agent Skills (arXiv:2608.14036, August 14) uses controlled experiments to answer "when are skills useful, why, and where do they stop working"—the perfect empirical companion to this site's Skills page; AgentRewind (arXiv:2608.14380, August 14) tackles recoverability in long-horizon execution—early errors contaminate both the context and the environment state, later actions usually can't undo them, so execution has to be rollback-capable; HarnessRisk (arXiv:2608.17597, August 18) is the first lifecycle safety benchmark organized by harness responsibility (tools, extensions, persistent state, permissions, external actions), pushing permissions and safety from best practice into the measurable.
Evaluation: From Scores to Discipline
Evaluation research after 2025 runs on two parallel tracks: harder new benchmarks, and a systematic interrogation of whether the scores themselves can be trusted.
The new benchmarks fill the old benchmarks' blind spots:
| Benchmark | Year | What it measures | Key finding / characteristics |
|---|---|---|---|
| τ²-bench (arXiv:2506.07982) | 2025 | Dual control: user and agent can both operate the shared environment (customer-service scenarios) | Agents degrade sharply when the user can also act; pass^k consistency across repeated runs is low |
| BrowseComp (arXiv:2504.12516) | 2025 | Answers findable only through persistent multi-hop browsing, 1,266 hard questions | Ordinary models nearly all wipe out; it triggered the arms race in deep-research-style agents |
| Terminal-Bench | 2025 | Tasks in a real terminal environment (compiling, configuring, ops) | Pushes evaluation into the coding agent's home turf—the command line |
Evaluation discipline became a subject of its own. Princeton's HAL (Holistic Agent Leaderboard) (2025) turned evaluation into infrastructure: run multiple benchmarks on one unified harness, publish full trajectories, disclose costs, so anyone can audit "how was this score produced". The Confucius Code Agent (arXiv:2512.10398, December 2025) quantified the harness's contribution with controlled experiments comparing scaffolds on the same model; May 2026's Stop Comparing LLM Agents Without Disclosing the Harness (arXiv:2605.23950) drew the hard line: once re-tested under standardized scaffolds, agent scores without harness disclosure are not comparable.
This site's reading advice
For any agent paper or leaderboard published after 2025, run three checks first: Is the harness disclosed? Is it a single run or pass^k? Are failure cases analyzed? Miss two of the three and the numbers are advertising copy. The same discipline applies when you self-test a harness you've built—see design principles.
The Seven Threads Taken Together
- Long horizons tell us the bottleneck is execution reliability and error compounding, not per-step cleverness;
- Context engineering and memory answer "how information gets in, and how it stays";
- The multi-agent cold water reminds us that parallelization gains must outrun coordination costs;
- Self-improvement turns the harness into an optimization target and raises the standing of verification infrastructure;
- Endogenous training goes a step further: the harness moves from deployment shell to training environment, and the skeleton's engineering quality starts to shape model capability directly;
- Evaluation discipline supplies the test for all of the above.
One observation runs through it all: papers from 2022–2024 asked "what can the model do"; papers from 2025–2026 ask "what can the system verifiably do, over what time scale, at what cost". The center of gravity moved from capability showcase to engineering accountability—that in itself is the harness perspective winning.
Further Reading
- Core papers, close up—the predecessors of everything on this page: ReAct, Reflexion, MemGPT, SWE-agent
- Paper map—the complete academic map, organized by theme
- What is an Agent Harness—a detailed read of Confucius CCA and the harness-disclosure manifesto
- Context engineering—the engineering treatment of Context Rot, recitation, and compression
- Memory systems—the product form of the A-MEM / Mem0 / MemOS ideas
- Subagents—the engineering default amid the multi-agent hype
- Planning—how externalized state counters multi-turn drift, at the mechanism level
- Skills—where a self-improving agent's skill consolidation lands in products
- Observability—verification and audit infrastructure for the self-improvement era
- Common pitfalls—engineering lessons the frontier papers keep confirming
References
- Measuring AI Ability to Complete Long Tasks (arXiv:2503.14499)
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs (arXiv:2509.09677)
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents (arXiv:2502.15840)
- Anthropic: Project Vend (Claude runs a physical vending machine)
- Manus blog: Context Engineering for AI Agents: Lessons from Building Manus
- Anthropic engineering blog: Effective Context Engineering for AI Agents
- Chroma research report: Context Rot
- LLMs Get Lost in Multi-Turn Conversation (arXiv:2505.06120)
- Recursive Language Models (arXiv:2512.24601)
- A-MEM: Agentic Memory for LLM Agents (arXiv:2502.12110)
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory (arXiv:2504.19413)
- MemOS: An Operating System for Memory-Augmented Generation (arXiv:2505.22101)
- Anthropic engineering blog: How we built our multi-agent research system
- Why Do Multi-Agent LLM Systems Fail? (MAST, arXiv:2503.13657)
- A Self-Improving Coding Agent (SICA, arXiv:2504.15228)
- Darwin Gödel Machine (arXiv:2505.22954)
- Agentic Context Engineering (ACE, arXiv:2510.04618)
- A Survey of Self-Evolving Agents (arXiv:2507.21046)
- τ²-bench: Evaluating Conversational Agents in a Dual-Control Environment (arXiv:2506.07982)
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents (arXiv:2504.12516)
- Terminal-Bench (Laude Institute)
- HAL: Holistic Agent Leaderboard (Princeton)
- Confucius Code Agent (arXiv:2512.10398)
- Stop Comparing LLM Agents Without Disclosing the Harness (arXiv:2605.23950)
- Agent Lightning v1.0: Towards Harnessed Agentic RL (arXiv:2608.17528)
- LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents (arXiv:2608.17393)
- ClawGym II: Exploring Black-Box RL on Agent Harness (arXiv:2608.16798)
- Demystifying Agent Skills: Why They Work—Until They Don't (arXiv:2608.14036)
- AgentRewind: Recoverable Execution for Long-Horizon LLM Agents (arXiv:2608.14380)
- HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety (arXiv:2608.17597)