Skip to content

Evaluation & Observability

At a glance Why agent evals are hard, trace recording and replay, the major benchmarks (SWE-bench, Terminal-Bench, τ-bench, WebArena/WorkArena), how to use LLM-as-judge and where its biases lie, and how to build an evals-driven harness iteration workflow.

Evaluation & Observability ​

Every design decision in an agent harness—how context gets assembled, how tools are described, when to plan, when to ask for help—ultimately has to answer the same question: after the change, is the agent actually better? Without reliable evals and observability, harness engineering is flying blind: you tweak the system prompt and it feels smarter, but in reality you may just have gotten lucky last run.

This page covers three things: how to measure agent behavior (evals), how to see agent behavior clearly (observability), and how to fuse the two into an iteration flywheel.

Why agent evals are hard ​

Every assumption behind traditional software testing breaks down in the face of agents:

Non-determinism. Run the same task on the same harness twice and you may get different results. Temperature, sampling, the timing of tool responses, even the model provider's backend load balancing all inject variance. A single run's pass/fail carries almost no information—you have to care about distributions. The pass^k metric introduced by τ-bench (the share of tasks where the agent succeeds on all of k independent runs) quantifies this brutally: an agent at pass@1 of 90% is down to roughly 43% at pass^8 (0.9⁸ ≈ 0.43). "Can do it" and "does it reliably" are two different capability tiers.

Long trajectories. A single agent run involves dozens to hundreds of model calls and tool interactions. When the final result is right, the process may have been dumb luck; when it's wrong, 90% of the steps may have been fine, with one wrong parameter at the very last step. Looking only at the final result, you can't tell which part of the harness to change.

Environment dependence. An agent's side effects land in real environments: file systems, Docker containers, browsers, live APIs. Evals need reproducible environment snapshots, and the environments themselves drift—SWE-bench tasks depend on specific versions of open-source repos, WebArena requires self-hosting an entire suite of websites, and every Terminal-Bench task runs in its own Docker container. Environment setup costs far more than "prepping a batch of prompts."

Ambiguous success criteria. For "fix this issue," success means all tests pass—but the tests themselves may be incomplete. For "rebook the user's flight," success means the final database state complies with policy—whether the user is happy is another matter. τ²-bench's dual-control environment (both the agent and a simulated user can change the shared state) means even the "final environment state" isn't something the agent alone can determine.

A common self-deception

Twenty tasks, one run each, scored by "this response looks pretty good to me"—that's not an eval, that's divination. Small sample + single-shot sampling + subjective judgment: stack three noise sources together and your conclusions are essentially random.

Traces: recording and replay ​

The raw material for evaluation is the trace (trajectory): a complete, structured record of one agent run. A minimal viable trace schema:

json
{
  "run_id": "run_01H...",
  "task": { "id": "swe-verified__django-16379", "input": "..." },
  "config": {
    "model": "claude-opus-4.6",
    "harness_version": "git:a3f9c2",
    "system_prompt_hash": "sha256:...",
    "tools": ["read", "edit", "bash"]
  },
  "steps": [
    {
      "step": 7,
      "type": "tool_call",
      "llm_input_tokens": 48210,
      "llm_output_tokens": 312,
      "latency_ms": 3400,
      "tool": { "name": "bash", "args": {"cmd": "pytest -x"}, "result": "..." }
    }
  ],
  "outcome": { "status": "resolved", "cost_usd": 1.87, "wall_time_s": 623 }
}

Three design points:

  1. Record the harness config, not just the model. Swap the system prompt and tool descriptions around the same model and scores can move by several points—the Holistic Agent Leaderboard research found that the same Claude Opus 4 on GAIA dropped from 64.9% to 57.6% when paired with a different agent framework. Without the harness version in the trace, attribution after the fact is impossible.
  2. Log token counts and latency at every step. This is the only data source for cost analysis and context-engineering optimization (see Context Engineering).
  3. Traces must be replayable. Freeze the model outputs, swap out the prompt or tool result at one step, and you can run counterfactual experiments: "what if the context had been assembled this way?" This is the microscope of harness debugging.

You don't have to reinvent the wheel: OpenTelemetry's GenAI semantic conventions already standardize the gen_ai.* attribute family—gen_ai.system, gen_ai.request.model, gen_ai.usage.input_tokens, the span structure for tool calls, and so on. Have your harness emit OTLP spans directly and any compatible backend can sit downstream, so you avoid lock-in to a single platform.

text
┌───────────────────── Agent Run (trace) ──────────────────────┐
│  span: llm.call#1   model=opus-4.6   in=12k out=0.4k 2.1s    │
│  span: tool.read    path=src/auth.py              ok 12ms    │
│  span: llm.call#2   model=opus-4.6   in=31k out=1.2k 4.8s    │
│  span: tool.bash    cmd="pytest -x"          FAIL 8.2s       │
│  span: llm.call#3   model=opus-4.6   in=48k out=2.0k 6.4s ←─ ┼─ failure point
│  span: tool.edit    path=src/auth.py              ok 9ms     │
│  ...                                                         │
│  outcome: resolved=false  cost=$1.87  steps=34               │
└──────────────────────────────────────────────────────────────┘

The major benchmarks: what each one measures ​

Before picking a benchmark, be clear about which agent capability it measures. The four most-cited ones:

BenchmarkWhat it measuresEnvironmentSuccess criteriaCurrent state (mid-2026)
SWE-bench Verified (500 tasks)Fixing real GitHub issues: read the repo, locate the bug, write a patchReal Python repos + test suitesPatch passes the FAIL_TO_PASS testsTop scores above 80%, nearing saturation, contamination disputes growing
Terminal-Bench 2.0 (89 tasks)Real terminal work: compiling, configuring servers, data processing, security tasksIndependent Docker container per task, with human oracle solutions and final-state testsContainer's final state passes the testsDesigned at launch to hold frontier models below 50%; as of August 2026 the top score is ~83%, the average ~58%
τ-bench / τ²-benchFollowing policies and calling tool APIs correctly across multi-turn conversationsLLM-simulated user + programmable domain APIs (retail/airline/telecom)Final database state matches the annotated statepass^1 can reach 80%+, but pass^k consistency falls off fast; τ² telecom once pushed GPT-4.1 down to 34%
WebArena / WorkArenaCompleting tasks on real web UIsSelf-hosted website clusters (e-commerce/GitLab/forums/maps); WorkArena runs on ServiceNowProgrammatic checks of page/backend stateOriginal WebArena paper: GPT-4 scored just 14.4% (humans 78.2%); GPT-4o at 42.7% on WorkArena, climbing fast in recent years

A few takes:

  • SWE-bench Verified is losing its discriminating power. From a 1.96% baseline in October 2023 to 80%+ in 2026, its historical mission is nearly complete. More noteworthy are its derivatives (Multimodal, Multilingual, and the harder SWE-bench Pro—where the best public-set score is only around 44%). The community generally discounts high Verified scores: concerns about contamination (having seen these issues in training data) and overfitting to the leaderboard persist, and OpenAI has stopped submitting results to Verified.
  • Terminal-Bench is currently the leaderboard most sensitive to harness quality. It explicitly scores "agent + model" as a single system, and submissions must declare their harness. The LangChain team published a textbook example: without changing the model, just by changing the harness (adding a pre-completion checklist, loop detection, and other middleware), Terminal-Bench 2.0 went from 52.8 to 66.5, jumping from outside the top 30 into the top 5.
  • The τ-bench family's core contribution isn't the scores—it's the pass^k metric. It forces you to confront consistency rather than one-off luck, which is the only thing production systems care about.
  • The value of WebArena-style benchmarks is grounding. The model has to translate natural-language intent into specific buttons and forms, which text-only benchmarks can't test. Its pain points are heavy environment operations—you maintain the self-hosted website clusters yourself—and results that are extremely sensitive to how the harness represents the browser (AXTree vs. screenshots vs. hybrid).

Practical advice on choosing benchmarks

For coding harnesses use Terminal-Bench or the SWE-bench family; for conversational/customer-support agents use τ²-bench; for browser agents use WebArena/WorkArena (via BrowserGym as the unified interface). Always report the mean and variance across multiple runs—better yet, report pass^k directly.

Outcome-based vs. trajectory-based evaluation ​

Outcome-based evaluation looks only at the final state: did the patch pass the tests, is the database state right, did the task succeed. The upside: it's objective, cheap, and automatable at scale. The downside: it can't answer "why," and reward hacking is hard to fully guard against—an agent can "pass" an SWE-bench task by deleting the tests.

Trajectory-based evaluation inspects the process: whether each tool choice was sensible, whether the agent spun in circles, whether it hallucinated APIs that don't exist, whether there were files it should have read but didn't. It can pinpoint specific harness defects, but it's hard to automate—"was this step reasonable?" is itself a judgment call.

In practice the division of labor is: outcome-based evals act as the regression gate (full run on every harness change), trajectory-based evals do failure attribution (sample and analyze the failures). A useful middle layer is trace heuristics—programmatic rules that need no LLM judgment:

python
def trace_heuristics(trace):
    issues = []
    if consecutive_identical_tool_calls(trace) >= 3:
        issues.append("loop_detected")          # spinning in circles
    if any(s["llm_input_tokens"] > 0.8 * MAX_CTX for s in trace.steps):
        issues.append("context_near_overflow")  # context near overflow
    if trace.outcome.status == "resolved" and trace.outcome.cost_usd > 5:
        issues.append("success_but_wasteful")   # successful but expensive
    if final_answer_cites_unread_files(trace):
        issues.append("grounding_violation")    # cites files it never read
    return issues

These rules run on every trace at nearly zero cost, yet they catch 80% of the common failure patterns during harness iteration.

LLM-as-judge: usage and biases ​

For open-ended tasks (writing summaries, editing copy, open-ended QA), there is no programmatic judge—you have no choice but to let another LLM play judge. The right way to run LLM-as-judge:

  • Anchor scoring to a rubric. "Rate it 1-5" is inferior to "check whether it meets these 4 criteria, 0/1 each." Per-criterion binary judgments are far more stable than holistic scores.
  • Use pairwise comparison instead of absolute scores. "Which is better, A or B?" has a higher signal-to-noise ratio than "how many points is A worth," especially for comparing two harness versions.
  • Give the judge the trace, not just the final answer. Asking "did this agent do well?" without showing it the intermediate steps is like doing code review with nothing but the commit message.

But the biases of LLM-as-judge are solidly documented in the literature (Zheng et al., 2023—the MT-Bench/Chatbot Arena paper measured them systematically):

BiasHow it shows upMitigation
Position biasIn pairwise judgment, favors the answer shown firstJudge twice with the order swapped; flag inconsistencies
Verbosity biasPrefers longer answers regardless of qualityExplicitly penalize verbosity in the rubric
Self-preferenceA GPT-family judge scores GPT-family agents higherUse a model from a different vendor for the judge
Score compressionScores bunch up at 3.5-4.5 with poor discriminationBinary rubric + pairwise

WARNING

LLM-as-judge is suited to relative comparison and coarse filtering, not as the sole basis for a release gate. Any conclusion along the lines of "the judge says the new version is 2 points better" should be spot-checked by human review of a 5-10% sample before you trust it.

Production monitoring: cost, latency, success rate ​

Benchmarks govern "capability"; once you ship, you also have to govern "operations." Agent products have monitoring metrics that differ from traditional APIs. Four core groups:

Metric groupSpecific metricsAlert signals
Success rateTask completion rate, pass^k (periodic re-runs on critical tasks), human intervention rateWeek-over-week drop >5%
CostTokens per task, dollars per task, cost per successful task (= total cost / successes)Success rate flat but costs rising—the harness is spinning its wheels
LatencyEnd-to-end wall-clock time, per-step LLM latency p50/p95, tool duration distributionDeteriorating p95 tail latency, usually from context bloat
HealthAverage step count, loop-detection trigger rate, context utilization, permission denial rate (see Permissions & Human-in-the-Loop)A rising average step count often precedes a falling success rate

Watch the "cost per successful task" composite metric—it's far more honest than cost or success rate alone. The first symptom of a broken harness change is often not a falling success rate but "paying more to get the same thing done."

The evals-driven harness iteration workflow ​

Assemble all the parts above into a closed loop:

text
┌────────────────────────────────────────────────────────────┐
│                  evals-driven iteration loop               │
│                                                            │
│   ① Collect failures ◄──── production traces / human labels │
│   ② Freeze into eval cases (task + env snapshot + judge)   │
│   ③ Change the harness (prompts/tools/context strategy)    │
│   ④ Run full eval suite (N runs/task; mean+variance+pass^k)│
│   ⑤ Trace attribution: new failures vs. old—new way to die?│
│   ⑥ Ship if targets met, archive traces ──────► back to ①  │
└────────────────────────────────────────────────────────────┘

A few hard-won disciplines from practice:

  1. Grow the eval set out of real failures; don't invent test cases in a vacuum. Failure cases mined from production traces are an order of magnitude more valuable than test questions dreamed up at a desk. Every time a user reports a bad case, the first reflex should be "turn it into an eval."
  2. Judge first. When adding an eval case, write the judge first (tests, state checks, rubric) and confirm it actually scores a known failure as a failure—then talk about fixing the harness. The judge itself needs testing too.
  3. Guard against overfitting: keep a holdout. The eval set you stare at during iteration slowly gets "taught to the test"—the harness tunes itself to those cases, scores climb, generalization doesn't. Keep a holdout set whose details you never inspect, run only at milestones.
  4. Change one variable at a time. Swap the model, rewrite the prompt, and add a tool all at once, and when the score moves you won't know whom to credit. The harness config hashes recorded in traces exist to support exactly this kind of controlled experiment.
  5. Report effect sizes, not single-run scores. "From 61% to 63%" is noise at N=100 with single-run sampling; only as a 5-run mean plus a paired test might it be signal.

The most disciplined form of this process is what LangChain did on Terminal-Bench: treat harness changes as testable experiments, use middleware (pre-completion self-checks, loop detection)—pluggable harness components—as the variables, and verify the gains one at a time. This echoes the claim this site keeps making: the model vs. harness divide determines where your main optimization battlefield is.

Tooling ecosystem ​

ToolPositioningNotes
OpenTelemetry GenAI semantic conventionsFoundational standardgen_ai.* attributes and span model, vendor-neutral, the default foundation for harness instrumentation
LangfuseOpen-source observability platformSelf-hostable, trace/score/eval in one, MIT-style license, a good fit for data-sensitive settings
LangSmithLangChain ecosystem platformDeep integration with LangGraph/LangChain, excellent dataset + replay experience, commercial SaaS
Braintrust / Arize Phoenix / WeaveEval platformsEach with its own emphasis: Braintrust leans into eval workflows, Phoenix is open source and trace-analysis focused, Weave belongs to the W&B ecosystem
BrowserGym / AgentLabWeb agent eval environmentsUnified access to WebArena/WorkArena/MiniWoB and more, built by ServiceNow
HarborTerminal agent eval runtimeThe official Terminal-Bench runtime, also runs custom containerized tasks

The selection advice is plain: instrument with the OTel standard (no lock-in), choose self-hosted vs. SaaS platforms based on data-compliance requirements, and use the benchmark's official eval runtime (rewriting the judging logic yourself makes your numbers incomparable with the leaderboard—the most common way teams shoot themselves in the foot).

Trade-offs ​

  • Eval coverage vs. iteration speed. If a full eval suite takes 8 hours and burns $500 per run, engineers will route around it. Tier it: smoke (20 tasks/5 minutes/every commit), regression (500 tasks/1 hour/every merge), full (everything + multi-sampling/at milestones).
  • Judge cost vs. judgment quality. Programmatic judges are expensive to write and free to run; LLM judges are quick to write, billed per token at run time, and carry biases. Anything that can be judged programmatically, judge programmatically—always.
  • Observability granularity vs. overhead. Persisting full inputs and outputs for every tool call makes storage and retrieval costs nontrivial on long trajectories. The usual compromise: full span metadata for everything + sampled retention for large payloads.
  • Leaderboard comparability vs. business relevance. Public benchmarks guarantee comparability, but your business distribution is worlds apart from SWE-bench's. Use public leaderboards to pick models and calibrate; a private eval set is what drives day-to-day iteration—neither can replace the other.
Further reading

References ​