Skip to content

Evaluation Systems

At a glance A systematic treatment of AI agent evaluation: why it's an order of magnitude harder than model evaluation, the four-layer framework from unit to trajectory to outcome to system, LLM-as-judge biases and calibration, the current landscape of SWE-bench, GAIA, OSWorld, τ-bench and other benchmarks, and how to build a contamination-resistant private eval.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Evaluation Systems ​

If the prompt sets an agent's ceiling, then evals determine whether you can know how far you are from it. After 2025, one industry consensus solidified: model capability is no longer the bottleneck for most agent projects — evaluation capability is. Teams without evals spin in the loop of "tweak the prompt → it feels better → ship → blow up → roll back"; teams with evals turn every change into a comparable experiment. Anthropic's engineering blog puts it bluntly in "Demystifying Evals for AI Agents": without evals, teams fall into reactive mode — fixing problems only after production incidents, where each fix creates three more.

This page covers four things: why agent evaluation is hard, what to evaluate (four layers), what to evaluate with (four kinds of grader), and what the mainstream benchmarks actually measure. It closes with a methodology for building your own private eval and the score-gaming traps you must watch for.

1. Why Agent Evaluation Is Hard ​

Traditional NLP evaluation is a static game of "give an input, check the answer." Agent evaluation is nothing of the sort. The difficulty concentrates in four places.

1. Non-determinism: the same input runs differently every time ​

LLM sampling is inherently stochastic, and every one of the agent's decisions amplifies it: pick a different tool at step one and the entire trajectory diverges from there. This means "ran it once, it passed" tells you nothing. The industry-standard answer is pass^k (a metric introduced by τ-bench): run the same task k times; it's only truly reliable if all k pass. An agent with a pass@1 of 70% may score single digits on pass^8 — a life-or-death line for support and finance scenarios.

2. Multi-step trajectories: right and wrong are no longer binary ​

On a 20-step task, the agent's final answer is correct, but it mis-called tools three times in the middle and lucked out — does that count as success? Conversely, it failed at the end, but the cause was a website crashing at step 18 — is that the agent's fault? Outcome-only evaluation gets both of these wrong, which is why trajectory-level evaluation is necessary (see Section 2).

3. Environment dependence: what you're testing is "agent + environment" as a whole ​

For environment-style benchmarks like WebArena and OSWorld, task success depends simultaneously on model capability, scaffold quality, and environment stability. The same model on a different scaffold can produce wildly different scores — on Princeton's HAL leaderboard, Claude Sonnet 4.5 scored 74.55% on GAIA under HAL's own Generalist scaffold but only 30.91% under HuggingFace's Open Deep Research scaffold, a 43-point gap coming entirely from the orchestration layer. So whenever you see a benchmark number, the first question to ask is: which scaffold?

4. Cost: one eval run may be far more expensive than you think ​

An agent eval isn't a matter of running a few thousand prompts. Terminal-Bench's official leaderboard annotates the API cost of each full evaluation run: Claude Code + Fable 5 costs about $553 per run, and Codex + GPT-5.5 costs over $2,000 per run. That's before environment setup, failure retries, and human spot checks. Eval design must answer "how much money buys how much statistical confidence" — the evaluation-side mirror of the question discussed on the cost & optimization page.

A common misjudgment

Treating "the model's score on public benchmarks" as "my agent's success rate in production." The former is measured in a controlled environment, on a fixed scaffold, over a known task distribution; the latter faces your private data, dirty inputs, and long-tail requests. The gap between them is often large enough to make benchmark rankings useless as a reference — that's exactly what Section 7 of this page unpacks.

2. The Four Layers of Evaluation ​

A mature agent evaluation system is layered, like unit tests → integration tests → end-to-end tests → production monitoring in software. Each layer answers a different question, and mixing them distorts conclusions.

┌─────────────────────────────────────────────────────────────┐
│ L4 System     cost / latency / safety / user satisfaction   │  ← is it worth shipping
├─────────────────────────────────────────────────────────────┤
│ L3 Outcome    task success rate (end-to-end, env-verified)  │  ← did the work get done
├─────────────────────────────────────────────────────────────┤
│ L2 Trajectory per-step decision quality, tool choice,       │  ← was the process right
│               error recovery                                │
├─────────────────────────────────────────────────────────────┤
│ L1 Unit       a single tool call, a single prompt's output  │  ← is the part any good
└─────────────────────────────────────────────────────────────┘

L1: the unit layer ​

Take the agent apart and evaluate the pieces: is a single tool's schema filled correctly, what's the classification accuracy of a single prompt, what's the recall of the RAG retrieval. This layer is closest to traditional LLM evaluation and can run fast against static datasets + assertions. Its value is problem localization — when end-to-end fails, you need to know which part broke. The mature reference point for tool calling at this layer is the Berkeley Function-Calling Leaderboard (BFCL), which uses AST matching to judge function selection and argument filling.

L2: the trajectory layer ​

Evaluate the "process" rather than the "result": did the agent choose a sensible tool sequence, are there pointless loops, does it recover from errors, is the step count under control. The common approach: record the full trace (see Observability), then score trajectories with rules (e.g., "the same tool must not fail 3 times in a row") or LLM-as-judge. The trajectory layer's value is that it still finds problems when the outcome succeeded by luck.

L3: the outcome layer ​

End-to-end task success: given the task, is the final state right? SWE-bench's "tests pass = solved" and τ-bench's "database final state matches" both belong here. This is the most convincing layer and also the most expensive — it requires an executable environment or a reliable judge. When reading agent papers, this is the only layer of metric worth taking seriously (see how to read benchmark papers in the paper reading paths).

L4: the system layer ​

The agent as a live system's overall metrics: cost per task, P50/P99 latency, safety incident rate (see Security & Alignment), user adoption/satisfaction. No benchmark exists for this layer; it can only come from production instrumentation and A/B experiments. One counterintuitive piece of experience: L3 improving while L4 degrades is the norm — more accurate models are usually slower and more expensive, and whether that's worth it is decided by business metrics.

The first principle of layering

Lower layers are for debugging; upper layers are for decisions. L1 green but L3 red means the problem is in orchestration; L3 green but L4 red means the direction itself is wrong. Don't report "the agent got better" to your boss based on unit-layer improvements.

3. The Four Types of Grader ​

Once you've settled what to evaluate, the next question is what to evaluate with. In practice, the four grader types form a spectrum ordered by trustworthiness and cost, and mature eval systems combine them.

1. Rule assertions (code-based graders) ​

Regex matching, JSON schema validation, state assertions ("the order status in the database must become refunded"). Pros: deterministic, cheap, reproducible. Cons: they can only judge structured outputs with explicit criteria. Wherever a rule works, use a rule — the first principle of every mature team.

2. Environment verification (execution-based graders) ​

The reinforced version of rule assertions: instead of checking output text, check the state of the world. SWE-bench runs the test suite, OSWorld executes evaluation scripts to inspect file contents, Terminal-Bench does if-and-only-if verification inside containers — all in this class. It's the gold standard of outcome-layer evaluation, because it cannot be fooled by "eloquence" — the code either passes the tests or it doesn't. The cost: building and maintaining the environment is extremely expensive.

3. Model scoring (LLM-as-judge) ​

Have another LLM play judge and evaluate open-ended outputs (summary quality, conversation appropriateness, trajectory plausibility). This is currently the widest-coverage category — and the easiest to misuse.

The foundational work is Zheng et al.'s 2023 "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (NeurIPS 2023, arXiv:2306.05685), which named the bias taxonomy still used today:

  • Position bias: in pairwise comparison, the judge favors the answer placed first (or last). Mitigation: swap the order, judge twice, average.
  • Verbosity bias: longer answers win more often, regardless of quality. Mitigation: state explicitly in the prompt that "length is not a quality signal," or normalize for length.
  • Self-enhancement / self-preference bias: the judge favors content produced by its own model family. Panickssery et al.'s 2024 NeurIPS paper (arXiv:2404.13076) confirmed that LLMs can recognize and prefer their own outputs. Mitigation: use a different vendor's model family for the judge than for the model under test.

The same paper provides the positive evidence: GPT-4 judges agree with human preferences over 80% of the time, approaching human-vs-human agreement — the fundamental reason LLM-as-judge is engineering-viable. But note the boundary: the result holds for pairwise comparisons with a clear better/worse. For fine-grained scoring that requires deep domain knowledge (like "is this code's concurrency handling safe"), judge error rates rise significantly.

Calibration is a required course: keep a small batch of human-labeled samples (even just 50-100) and periodically check the judge's agreement with human judgment; below threshold, switch judges, revise the rubric, or fall back to humans. LangChain and other teams' 2026 practice guides list "continuously calibrating the judge with human-corrected samples" as a core process. Remember: the judge model is itself a system under test, and its scores are not evidence until validated.

4. Human spot checks (human review) ​

The most expensive and the most irreplaceable. Three uses: calibrating LLM judges, reviewing dimensions benchmarks can't cover (user experience, business compliance), and root-causing failure cases. The realistic allocation: rules and environment verification cover everything they can, LLM judges handle open-ended outputs, and humans spot-check 5%-10% of judge results.

python
# A pairwise-comparison judge with position-bias mitigation (sketch)
# Principle: evaluate twice with swapped order; adopt the verdict only if both agree

import random

JUDGE_PROMPT = """You are evaluating two responses from a customer support agent.
Criteria (in priority order):
1. Did it solve the user's problem  2. Did it comply with the refund policy  3. Is the tone professional
Note: response length is not a quality signal.

User question: {query}
Response A: {response_a}
Response B: {response_b}

Output JSON only: {{"winner": "A" | "B" | "tie", "reason": "one-sentence rationale"}}"""

def judge_pair(query, resp1, resp2, judge_fn):
    """judge_fn: a function that calls the judge model. Evaluate once per ordering."""
    if random.random() < 0.5:
        resp1, resp2 = resp2, resp1  # randomize which goes first, to avoid a fixed starting point
    votes = []
    for a, b in [(resp1, resp2), (resp2, resp1)]:  # swap positions
        result = judge_fn(JUDGE_PROMPT.format(query=query, response_a=a, response_b=b))
        votes.append(result["winner"])
    # map both votes back to the original responses; declare a winner only if they agree, else tie
    if votes[0] == "A" and votes[1] == "B":
        return "resp1"
    if votes[0] == "B" and votes[1] == "A":
        return "resp2"
    return "tie"  # position-inconsistent = the judge itself is unsure; escalate to a human

4. A Tour of the Mainstream Benchmarks ​

As of August 2026, agent benchmarks have moved from "is there one" to "which do you trust." First the overview, then each one's positioning and current state.

BenchmarkDomainSizeScoringStatus (2026-08)
SWE-bench Verifiedreal GitHub issue fixes500 taskstest suite passesSaturated; OpenAI announced its retirement in 2026-02
SWE-bench Prolong-horizon software engineering1,865 tasks (731 public)test suite passesOne of the hardest live coding benchmarks
GAIAgeneralist assistant tasks466 tasks (3 difficulty levels)exact answer matchWidely used, but hugely scaffold-dependent
WebArenaweb operation812 tasksprogrammatic verificationThe classic environment, increasingly carried forward by the Verified checker and successors
OSWorld(-Verified)real desktop OS operation369 tasksscript-based verification2.0 released 2026-06; the de facto standard for computer use
τ-bench / τ²-benchtool-user interaction, policy compliance3 domainsdatabase final state + LLM judgeThe benchmark for support scenarios; top scores nearing saturation
BrowseCompdeep web retrieval1,266 questionsshort-answer verificationThe core battleground for deep-research agents
Terminal-Benchterminal/CLI tasksmultiple versions (2.1 / 3.0)iff verification in containersActive leaderboard with cost data; the most engineering-flavored

SWE-bench (Verified / Pro) ​

SWE-bench (ICLR 2024) extracted 2,294 real GitHub issues from 12 popular Python repositories; the agent must produce a patch that turns fail-to-pass tests green without breaking pass-to-pass tests. It defined the engineering standard of the "environment verification" evaluation paradigm.

Verified is the human-curated subset OpenAI released in 2024 (500 tasks), removing samples with unclear descriptions or unreliable tests; for the following two years it was the most-cited number for coding agents. But by 2026 it had completed its historical mission: the official leaderboard's top crossed 76% at the end of 2025, and third-party trackers showed 75%+ and climbing through 2026 (aggregators differ wildly in methodology — verify before citing). More importantly, in February 2026 OpenAI published a post announcing it would no longer use SWE-bench Verified to evaluate its own models, citing increasing contamination and flawed tests — it "increasingly fails to measure differences at the coding frontier." A benchmark being publicly retired by a leading lab is an event worth remembering in evaluation history.

SWE-bench Pro (Scale AI, 2025-09, arXiv:2509.16941) is the successor: 1,865 tasks across 41 repositories, designed specifically against contamination — all 731 public tasks come from repos under strong copyleft licenses like GPL (legally blocking their entry into commercial training corpora), plus a private set (276 tasks from startups' proprietary code) and a fully secret held-out set (858 tasks). It is also genuinely harder: the reference patches average 107.4 lines changed across 4.1 files. At release, GPT-5 and Claude Opus 4.1 scored only 23.3% / 23.1% on the public set, dropping further to 14.9% / 17.8% on the private set — the fall from Verified's 70%+ to Pro's 23% is what "contamination + saturation" looks like when quantified. If you need to cite a coding-agent capability number in 2026, Pro is far more credible than Verified.

GAIA ​

GAIA (Meta, HuggingFace, and the AutoGPT team jointly; arXiv:2311.12983) measures the "generalist assistant": 466 questions that are conceptually simple for humans and maddening for agents, spanning web retrieval, file parsing, multimodality, and multi-step reasoning, graded into three difficulty levels by how long they take a human. Answers are short-text exact matches, so scoring is clean.

When using GAIA numbers, you must check the scaffold. Princeton's HAL leaderboard (now paused for new models, pivoting toward agent-reliability research) shows that the same batch of models can differ by tens of points between the HAL Generalist scaffold and HuggingFace's Open Deep Research scaffold — Claude Sonnet 4.5 at 74.55% vs 30.91%, Claude Opus 4.1 at 68.48% vs 28.48%. HAL also publishes the cost of each evaluation run (anywhere from tens to $2,800), making it one of the few leaderboards that puts the "cost-accuracy" Pareto front on the table.

WebArena and OSWorld ​

WebArena (ICLR 2024, arXiv:2307.13854) built an entire suite of reproducible simulated sites (Reddit, GitLab, e-commerce, maps, a wiki), with 812 long-horizon tasks judged by programmatic verification. At release, the GPT-4 agent's success rate was just 14.41% against 78.24% for humans — that gap defined the 2023-2024 web agent research agenda. Today it mostly lives on as the "WebArena Verified checker" inside successor work; fewer teams chase the original leaderboard directly.

OSWorld pushes the same idea onto real operating systems: 369 tasks run inside real Ubuntu/Windows/macOS virtual machines, involving cross-application workflows, each with an execution-based evaluation script. At release, humans completed 72.36% while the strongest model managed 12.24% — as brutal as WebArena in its day. It evolves fast: upgraded to OSWorld-Verified in 2025-07 (fixing community-reported broken samples, with AWS support compressing a full run to under an hour), then OSWorld 2.0 in 2026-06. Third-party trackers show the Verified leaderboard's top around 85% (not independently verified — check the official leaderboard before citing). For computer use, it's the current de facto standard.

The τ-bench family ​

Sierra Research's τ-bench has a unique insight: in enterprise support scenarios, completing the task doesn't matter — not violating policy does. An agent that books the right flight but skips the rebooking-fee rule scores zero, with no partial credit. Scoring relies on database final-state matching, and the metric is the pass^k mentioned earlier. τ²-bench (arXiv:2506.07982) upgraded to a "dual-control environment": the simulated user also holds tools (say, needing to toggle their own phone's airplane mode), so the agent must teach and guide the user through the operation, folding collaborative communication into the evaluation. The family kept evolving in 2026; top scores on older domains like airline/retail have been pushed near saturation (aggregators report the 99% range), and the new domains are where the differentiation lives.

BrowseComp ​

Released by OpenAI in April 2025 (arXiv:2504.12516), 1,266 questions specifically testing the ability to "dig one deliberately hidden piece of information out of the vast internet" — the answers are short and automatically verifiable, but finding them requires persistently paging through dozens of sites. It is the core battleground for deep-research products (OpenAI Deep Research, Gemini Deep Research, and their 2026 successors), and it spawned BrowseComp-ZH, a Chinese version. Third-party trackers showed top scores around 90% by August 2026 — the score inflation is visible to the naked eye, another curve of "single digits to saturation in two years."

Terminal-Bench ​

Agent evaluation in the terminal: tasks run in Docker containers, judged by strict all-or-nothing (if-and-only-if) verification, with the harbor runner reproducing everything in one command. Its official leaderboard is, in this writer's view, the most "engineer-friendly": every entry shows score ± confidence interval, scaffold, date, and the cost of a single evaluation run. In mid-2026 the 2.1 leaderboard's top sat around 83-84% (Claude Code + Fable 5 at 83.8%, Codex + GPT-5.5 at 83.1%), with the same model again showing visible gaps across CLI scaffolds. If you're building a command-line/DevOps agent, this is the most relevant public yardstick.

Quick rules for picking a benchmark

Prefer the one isomorphic to your scenario: for coding look at SWE-bench Pro, desktop operation at OSWorld, support policy compliance at τ-bench, retrieval research at BrowseComp, terminal ops at Terminal-Bench. Comparing rankings across scenarios is meaningless — they aren't measuring the same thing at all.

5. A Methodology for Building Your Own Eval ​

Public benchmarks answer "is this model any good"; private evals answer "is my agent any good on my business." The latter is where you should invest. A workable process:

1. Task-set design: grow it from failure logs, don't invent it in your head ​

  • Source priority: real production failure cases > internal dogfooding records > a product manager's hunches. The first two carry distributional truth for free.
  • Scale: 30-50 high-quality tasks beat 500 filler ones. Every task must have an explicit, machine-checkable success criterion — tasks you can't write criteria for don't go in the set.
  • Stratified sampling: stratify by task type, difficulty, and input length so every stratum is covered and the eval isn't dominated by one task type.
  • Mix in adversarial samples: deliberately construct dirty inputs, ambiguous instructions, and over-privileged requests (safety red-line cases are mandatory — see Security & Alignment).

2. Anti-contamination: the lifeline of a private eval ​

  • Tasks and answers never enter any training/fine-tuning pipeline, are not reused as prompt examples, and don't go into public repos.
  • Rotate 20%-30% of the questions periodically (say, quarterly); an eval whose questions never change gets "overfit" by the team without anyone noticing — everyone optimizes against those 50 questions, scores rise, capability doesn't.
  • If you use public benchmark data in a private eval, assume it's already contaminated: treat it as a regression test only, never as proof of capability.

3. Statistical significance: a score without error bars is noise ​

  • Non-deterministic agents must be sampled multiple times. Run every task at least 5-10 times and report means and confidence intervals, never the single best run.
  • Sample size determines resolution: on a 50-task eval, a 5-point "improvement" is most likely noise. A rough calculation: under a binomial distribution, ±1σ at n=50 is about 7 points. To resolve 2-3-point differences you need several hundred tasks, or you accept wider intervals.
  • Use paired comparisons to sharpen sensitivity: run old and new agent versions over the same tasks with the same random-seed budget, and compare paired differences rather than two independent scores — Miller's "Adding Error Bars to Evals" (arXiv:2411.00640) makes the case systematically: an eval is an experiment, and a score is a statistic with sampling error.
  • Report cost too: "+3% accuracy at 2.5× the per-task cost" isn't an improvement — it's a trade.

4. The rollout cadence ​

  1. Week 1: collect 30 real cases, hand-write the criteria, get the pipeline working with rule assertions + human scoring.
  2. Weeks 2-4: add an LLM judge for open-ended outputs, calibrated against 50 human-labeled samples.
  3. After that: evals into CI — every prompt change, model swap, or scaffold change triggers a run, with score deltas pasted into the PR description.
  4. Quarterly: rotate questions, re-check judge agreement, re-baseline costs.

6. Eval-Driven Development ​

With a private eval in place, the way you develop changes qualitatively: from "it feels better after my change" to "every change is a controlled experiment."

          ┌──────────────────┐
          │ Production failure logs │
          └────────┬─────────┘
                   ▼
       Failure attribution (read the traces)      ←── observability as the foundation
                   │
                   ▼
     New failure mode? ──yes──► Freeze it as an eval task (into the set first, then fix the bug)
                   │no
                   ▼
       Change prompt / tools / scaffold
                   │
                   ▼
       Run evals (multi-sampling + paired comparison)
                   │
         ┌─────────┴─────────┐
   significantly better      unchanged or worse
         │                      │
         ▼                      ▼
     canary rollout         roll back, read the traces, find the cause
         │
         ▼
   Confirm with production metrics (L4); new failures loop back to step one

Two key disciplines: first, into the set before the fix — a newly discovered failure mode gets written up as an eval task and confirmed to fail before you start fixing, which gives you a permanent regression test; second, every "optimization" passes the eval gate, including switching to a more expensive model. The full operational detail of this pipeline (tooling, CI integration, review checklists) unfolds on the Evals in Practice page; here we only set the frame. If you want to put this on a resume or in an interview answer, compare it against the "evaluation & reliability" row of the job-seeker knowledge map.

7. Watch Out: When Benchmark Scores Drift from Real Capability ​

Finally, a cold shower. The credibility crisis of 2025-2026 benchmarks is an openly discussed topic, and several structural causes are worth committing to memory by everyone who reads leaderboards:

  • Contamination is the norm, not the exception. SWE-bench went public at the end of 2023; models released since then have most likely seen those issues and fixes in their training corpora. SWE-bench Pro fighting back with copyleft repos and private sets is precisely the admission that the default assumption should be "already contaminated." When OpenAI itself announced Verified's retirement in February 2026, the industry signal couldn't have been clearer.
  • Scaffold is score. The 43-point GAIA gap above is not an outlier. Vendors publish leaderboard numbers using their own carefully tuned scaffolds, maximum reasoning effort, and generous retry budgets — hand you the same model's bare API and you won't reproduce them. A leaderboard measures the joint system of "model + scaffold + budget," while the marketing only mentions the model.
  • Saturation is outpacing benchmark iteration. SWE-bench Verified went from single digits in 2024 to 76%+ by the end of 2025 in barely more than a year; τ-bench's older domains and BrowseComp are tracing the same curve. The average lifespan of a benchmark is shrinking, and "topping the chart" carries more marketing value than information.
  • Leaderboard tasks ≠ your tasks. SWE-bench's issues are well-written by maintainers, discussed by the community, and covered by tests; what you face may be a one-line requirement, legacy code, and a repo with no tests. Distribution differences eat generalization for breakfast.

The pragmatic responses are only three: treat public leaderboards as initial screening, not a decision basis (when scores are close, pick the cheaper, faster one with the better ecosystem, not the #1); put the decision weight on your own private eval; and for any "topped the X leaderboard" claim, ask three questions first — which scaffold, what budget, and is there third-party reproduction.

The one-line summary

Public benchmarks tell you where the model's ceiling is; private evals tell you where your product's floor is. People building agents need to watch both, but only the latter is your moat.

References ​