Skip to content

Reading Discipline & FAQ

At a glance Common questions that come up when reading Agent Harness papers: whether success-rate numbers can be compared across papers, the conditions behind Reflexion's 91%, and how to read a benchmark paper.

Reading Discipline & FAQ ​

The most common trap when reading agent papers isn't failing to understand them. It's taking the numbers too seriously while skimming the mechanisms. This page answers the three questions that come up most often.

Can success rates be compared across papers? ​

No — at least not directly. The absolute numbers in these papers (success rates, pass@1) depend almost entirely on the model version and harness configuration of the moment: the same method, run on a different base model or given a different set of tool descriptions, can produce wildly different results. Cross-paper comparisons tell you very little.

What's worth taking away is the mechanism and the failure analysis — why the method works, and under what conditions it breaks — not the scores themselves. That's the reading discipline every route in Reading Paths emphasizes.

What's the story behind Reflexion's 91% on HumanEval? ​

The HumanEval pass@1 of 91% that Reflexion reports was achieved with iterative attempts and test feedback allowed. It's not on the same track as the 80% single-generation baseline.

The comparison itself is the clearest illustration of what a harness is worth: wrap the model in an outer loop that reflects on failure and retries, and performance improves by close to an order of magnitude. But when you cite this number, state the conditions — otherwise you're comparing an open-book result with a closed-book one. Source: Paper Map · Self-Improvement.

How should you read a benchmark paper? ​

Don't fixate on the leaderboard. Look at three things:

  1. Where the tasks come from (realism) — are they hand-written problems, synthetic ones, or real-world issues and live operating environments?
  2. How success is judged (execution-based or match-based) — execution-based judging (running the tests, inspecting environment state) is far more trustworthy than string matching.
  3. What the failures look like — failure modes tell you what a harness needs to compensate for, which is far more useful than a success rate.

Example: when the SWE-bench paper reported that "Claude 2 solved only 1.96%," what made that number significant wasn't that the model seemed dumb. It was proof that the bottleneck at the time lay in the interaction, not in the model's intelligence — a finding that directly motivated SWE-agent's research on interface design. Source: Paper Map · Evaluation Benchmarks.

Further Reading ​