Appearance
Evaluation & Observability
Every design decision in an agent harness—how context gets assembled, how tools are described, when to plan, when to ask for help—ultimately has to answer the same question: after the change, is the agent actually better? Without reliable evals and observability, harness engineering is flying blind: you tweak the system prompt and it feels smarter, but in reality you may just have gotten lucky last run.
This page covers three things: how to measure agent behavior (evals), how to see agent behavior clearly (observability), and how to fuse the two into an iteration flywheel.
Why agent evals are hard
Every assumption behind traditional software testing breaks down in the face of agents:
Non-determinism. Run the same task on the same harness twice and you may get different results. Temperature, sampling, the timing of tool responses, even the model provider's backend load balancing all inject variance. A single run's pass/fail carries almost no information—you have to care about distributions. The pass^k metric introduced by τ-bench (the share of tasks where the agent succeeds on all of k independent runs) quantifies this brutally: an agent at pass@1 of 90% is down to roughly 43% at pass^8 (0.9⁸ ≈ 0.43). "Can do it" and "does it reliably" are two different capability tiers.
Long trajectories. A single agent run involves dozens to hundreds of model calls and tool interactions. When the final result is right, the process may have been dumb luck; when it's wrong, 90% of the steps may have been fine, with one wrong parameter at the very last step. Looking only at the final result, you can't tell which part of the harness to change.
Environment dependence. An agent's side effects land in real environments: file systems, Docker containers, browsers, live APIs. Evals need reproducible environment snapshots, and the environments themselves drift—SWE-bench tasks depend on specific versions of open-source repos, WebArena requires self-hosting an entire suite of websites, and every Terminal-Bench task runs in its own Docker container. Environment setup costs far more than "prepping a batch of prompts."
Ambiguous success criteria. For "fix this issue," success means all tests pass—but the tests themselves may be incomplete. For "rebook the user's flight," success means the final database state complies with policy—whether the user is happy is another matter. τ²-bench's dual-control environment (both the agent and a simulated user can change the shared state) means even the "final environment state" isn't something the agent alone can determine.
A common self-deception
Twenty tasks, one run each, scored by "this response looks pretty good to me"—that's not an eval, that's divination. Small sample + single-shot sampling + subjective judgment: stack three noise sources together and your conclusions are essentially random.
Traces: recording and replay
The raw material for evaluation is the trace (trajectory): a complete, structured record of one agent run. A minimal viable trace schema:
json
{
"run_id": "run_01H...",
"task": { "id": "swe-verified__django-16379", "input": "..." },
"config": {
"model": "claude-opus-4.6",
"harness_version": "git:a3f9c2",
"system_prompt_hash": "sha256:...",
"tools": ["read", "edit", "bash"]
},
"steps": [
{
"step": 7,
"type": "tool_call",
"llm_input_tokens": 48210,
"llm_output_tokens": 312,
"latency_ms": 3400,
"tool": { "name": "bash", "args": {"cmd": "pytest -x"}, "result": "..." }
}
],
"outcome": { "status": "resolved", "cost_usd": 1.87, "wall_time_s": 623 }
}Three design points:
- Record the harness config, not just the model. Swap the system prompt and tool descriptions around the same model and scores can move by several points—the Holistic Agent Leaderboard research found that the same Claude Opus 4 on GAIA dropped from 64.9% to 57.6% when paired with a different agent framework. Without the harness version in the trace, attribution after the fact is impossible.
- Log token counts and latency at every step. This is the only data source for cost analysis and context-engineering optimization (see Context Engineering).
- Traces must be replayable. Freeze the model outputs, swap out the prompt or tool result at one step, and you can run counterfactual experiments: "what if the context had been assembled this way?" This is the microscope of harness debugging.
You don't have to reinvent the wheel: OpenTelemetry's GenAI semantic conventions already standardize the gen_ai.* attribute family—gen_ai.system, gen_ai.request.model, gen_ai.usage.input_tokens, the span structure for tool calls, and so on. Have your harness emit OTLP spans directly and any compatible backend can sit downstream, so you avoid lock-in to a single platform.
text
┌───────────────────── Agent Run (trace) ──────────────────────┐
│ span: llm.call#1 model=opus-4.6 in=12k out=0.4k 2.1s │
│ span: tool.read path=src/auth.py ok 12ms │
│ span: llm.call#2 model=opus-4.6 in=31k out=1.2k 4.8s │
│ span: tool.bash cmd="pytest -x" FAIL 8.2s │
│ span: llm.call#3 model=opus-4.6 in=48k out=2.0k 6.4s ←─ ┼─ failure point
│ span: tool.edit path=src/auth.py ok 9ms │
│ ... │
│ outcome: resolved=false cost=$1.87 steps=34 │
└──────────────────────────────────────────────────────────────┘The major benchmarks: what each one measures
Before picking a benchmark, be clear about which agent capability it measures. The four most-cited ones:
| Benchmark | What it measures | Environment | Success criteria | Current state (mid-2026) |
|---|---|---|---|---|
| SWE-bench Verified (500 tasks) | Fixing real GitHub issues: read the repo, locate the bug, write a patch | Real Python repos + test suites | Patch passes the FAIL_TO_PASS tests | Top scores above 80%, nearing saturation, contamination disputes growing |
| Terminal-Bench 2.0 (89 tasks) | Real terminal work: compiling, configuring servers, data processing, security tasks | Independent Docker container per task, with human oracle solutions and final-state tests | Container's final state passes the tests | Designed at launch to hold frontier models below 50%; as of August 2026 the top score is ~83%, the average ~58% |
| τ-bench / τ²-bench | Following policies and calling tool APIs correctly across multi-turn conversations | LLM-simulated user + programmable domain APIs (retail/airline/telecom) | Final database state matches the annotated state | pass^1 can reach 80%+, but pass^k consistency falls off fast; τ² telecom once pushed GPT-4.1 down to 34% |
| WebArena / WorkArena | Completing tasks on real web UIs | Self-hosted website clusters (e-commerce/GitLab/forums/maps); WorkArena runs on ServiceNow | Programmatic checks of page/backend state | Original WebArena paper: GPT-4 scored just 14.4% (humans 78.2%); GPT-4o at 42.7% on WorkArena, climbing fast in recent years |
A few takes:
- SWE-bench Verified is losing its discriminating power. From a 1.96% baseline in October 2023 to 80%+ in 2026, its historical mission is nearly complete. More noteworthy are its derivatives (Multimodal, Multilingual, and the harder SWE-bench Pro—where the best public-set score is only around 44%). The community generally discounts high Verified scores: concerns about contamination (having seen these issues in training data) and overfitting to the leaderboard persist, and OpenAI has stopped submitting results to Verified.
- Terminal-Bench is currently the leaderboard most sensitive to harness quality. It explicitly scores "agent + model" as a single system, and submissions must declare their harness. The LangChain team published a textbook example: without changing the model, just by changing the harness (adding a pre-completion checklist, loop detection, and other middleware), Terminal-Bench 2.0 went from 52.8 to 66.5, jumping from outside the top 30 into the top 5.
- The τ-bench family's core contribution isn't the scores—it's the pass^k metric. It forces you to confront consistency rather than one-off luck, which is the only thing production systems care about.
- The value of WebArena-style benchmarks is grounding. The model has to translate natural-language intent into specific buttons and forms, which text-only benchmarks can't test. Its pain points are heavy environment operations—you maintain the self-hosted website clusters yourself—and results that are extremely sensitive to how the harness represents the browser (AXTree vs. screenshots vs. hybrid).
Practical advice on choosing benchmarks
For coding harnesses use Terminal-Bench or the SWE-bench family; for conversational/customer-support agents use τ²-bench; for browser agents use WebArena/WorkArena (via BrowserGym as the unified interface). Always report the mean and variance across multiple runs—better yet, report pass^k directly.
Outcome-based vs. trajectory-based evaluation
Outcome-based evaluation looks only at the final state: did the patch pass the tests, is the database state right, did the task succeed. The upside: it's objective, cheap, and automatable at scale. The downside: it can't answer "why," and reward hacking is hard to fully guard against—an agent can "pass" an SWE-bench task by deleting the tests.
Trajectory-based evaluation inspects the process: whether each tool choice was sensible, whether the agent spun in circles, whether it hallucinated APIs that don't exist, whether there were files it should have read but didn't. It can pinpoint specific harness defects, but it's hard to automate—"was this step reasonable?" is itself a judgment call.
In practice the division of labor is: outcome-based evals act as the regression gate (full run on every harness change), trajectory-based evals do failure attribution (sample and analyze the failures). A useful middle layer is trace heuristics—programmatic rules that need no LLM judgment:
python
def trace_heuristics(trace):
issues = []
if consecutive_identical_tool_calls(trace) >= 3:
issues.append("loop_detected") # spinning in circles
if any(s["llm_input_tokens"] > 0.8 * MAX_CTX for s in trace.steps):
issues.append("context_near_overflow") # context near overflow
if trace.outcome.status == "resolved" and trace.outcome.cost_usd > 5:
issues.append("success_but_wasteful") # successful but expensive
if final_answer_cites_unread_files(trace):
issues.append("grounding_violation") # cites files it never read
return issuesThese rules run on every trace at nearly zero cost, yet they catch 80% of the common failure patterns during harness iteration.
LLM-as-judge: usage and biases
For open-ended tasks (writing summaries, editing copy, open-ended QA), there is no programmatic judge—you have no choice but to let another LLM play judge. The right way to run LLM-as-judge:
- Anchor scoring to a rubric. "Rate it 1-5" is inferior to "check whether it meets these 4 criteria, 0/1 each." Per-criterion binary judgments are far more stable than holistic scores.
- Use pairwise comparison instead of absolute scores. "Which is better, A or B?" has a higher signal-to-noise ratio than "how many points is A worth," especially for comparing two harness versions.
- Give the judge the trace, not just the final answer. Asking "did this agent do well?" without showing it the intermediate steps is like doing code review with nothing but the commit message.
But the biases of LLM-as-judge are solidly documented in the literature (Zheng et al., 2023—the MT-Bench/Chatbot Arena paper measured them systematically):
| Bias | How it shows up | Mitigation |
|---|---|---|
| Position bias | In pairwise judgment, favors the answer shown first | Judge twice with the order swapped; flag inconsistencies |
| Verbosity bias | Prefers longer answers regardless of quality | Explicitly penalize verbosity in the rubric |
| Self-preference | A GPT-family judge scores GPT-family agents higher | Use a model from a different vendor for the judge |
| Score compression | Scores bunch up at 3.5-4.5 with poor discrimination | Binary rubric + pairwise |
WARNING
LLM-as-judge is suited to relative comparison and coarse filtering, not as the sole basis for a release gate. Any conclusion along the lines of "the judge says the new version is 2 points better" should be spot-checked by human review of a 5-10% sample before you trust it.
Production monitoring: cost, latency, success rate
Benchmarks govern "capability"; once you ship, you also have to govern "operations." Agent products have monitoring metrics that differ from traditional APIs. Four core groups:
| Metric group | Specific metrics | Alert signals |
|---|---|---|
| Success rate | Task completion rate, pass^k (periodic re-runs on critical tasks), human intervention rate | Week-over-week drop >5% |
| Cost | Tokens per task, dollars per task, cost per successful task (= total cost / successes) | Success rate flat but costs rising—the harness is spinning its wheels |
| Latency | End-to-end wall-clock time, per-step LLM latency p50/p95, tool duration distribution | Deteriorating p95 tail latency, usually from context bloat |
| Health | Average step count, loop-detection trigger rate, context utilization, permission denial rate (see Permissions & Human-in-the-Loop) | A rising average step count often precedes a falling success rate |
Watch the "cost per successful task" composite metric—it's far more honest than cost or success rate alone. The first symptom of a broken harness change is often not a falling success rate but "paying more to get the same thing done."
The evals-driven harness iteration workflow
Assemble all the parts above into a closed loop:
text
┌────────────────────────────────────────────────────────────┐
│ evals-driven iteration loop │
│ │
│ ① Collect failures ◄──── production traces / human labels │
│ ② Freeze into eval cases (task + env snapshot + judge) │
│ ③ Change the harness (prompts/tools/context strategy) │
│ ④ Run full eval suite (N runs/task; mean+variance+pass^k)│
│ ⑤ Trace attribution: new failures vs. old—new way to die?│
│ ⑥ Ship if targets met, archive traces ──────► back to ① │
└────────────────────────────────────────────────────────────┘A few hard-won disciplines from practice:
- Grow the eval set out of real failures; don't invent test cases in a vacuum. Failure cases mined from production traces are an order of magnitude more valuable than test questions dreamed up at a desk. Every time a user reports a bad case, the first reflex should be "turn it into an eval."
- Judge first. When adding an eval case, write the judge first (tests, state checks, rubric) and confirm it actually scores a known failure as a failure—then talk about fixing the harness. The judge itself needs testing too.
- Guard against overfitting: keep a holdout. The eval set you stare at during iteration slowly gets "taught to the test"—the harness tunes itself to those cases, scores climb, generalization doesn't. Keep a holdout set whose details you never inspect, run only at milestones.
- Change one variable at a time. Swap the model, rewrite the prompt, and add a tool all at once, and when the score moves you won't know whom to credit. The harness config hashes recorded in traces exist to support exactly this kind of controlled experiment.
- Report effect sizes, not single-run scores. "From 61% to 63%" is noise at N=100 with single-run sampling; only as a 5-run mean plus a paired test might it be signal.
The most disciplined form of this process is what LangChain did on Terminal-Bench: treat harness changes as testable experiments, use middleware (pre-completion self-checks, loop detection)—pluggable harness components—as the variables, and verify the gains one at a time. This echoes the claim this site keeps making: the model vs. harness divide determines where your main optimization battlefield is.
Tooling ecosystem
| Tool | Positioning | Notes |
|---|---|---|
| OpenTelemetry GenAI semantic conventions | Foundational standard | gen_ai.* attributes and span model, vendor-neutral, the default foundation for harness instrumentation |
| Langfuse | Open-source observability platform | Self-hostable, trace/score/eval in one, MIT-style license, a good fit for data-sensitive settings |
| LangSmith | LangChain ecosystem platform | Deep integration with LangGraph/LangChain, excellent dataset + replay experience, commercial SaaS |
| Braintrust / Arize Phoenix / Weave | Eval platforms | Each with its own emphasis: Braintrust leans into eval workflows, Phoenix is open source and trace-analysis focused, Weave belongs to the W&B ecosystem |
| BrowserGym / AgentLab | Web agent eval environments | Unified access to WebArena/WorkArena/MiniWoB and more, built by ServiceNow |
| Harbor | Terminal agent eval runtime | The official Terminal-Bench runtime, also runs custom containerized tasks |
The selection advice is plain: instrument with the OTel standard (no lock-in), choose self-hosted vs. SaaS platforms based on data-compliance requirements, and use the benchmark's official eval runtime (rewriting the judging logic yourself makes your numbers incomparable with the leaderboard—the most common way teams shoot themselves in the foot).
Trade-offs
- Eval coverage vs. iteration speed. If a full eval suite takes 8 hours and burns $500 per run, engineers will route around it. Tier it: smoke (20 tasks/5 minutes/every commit), regression (500 tasks/1 hour/every merge), full (everything + multi-sampling/at milestones).
- Judge cost vs. judgment quality. Programmatic judges are expensive to write and free to run; LLM judges are quick to write, billed per token at run time, and carry biases. Anything that can be judged programmatically, judge programmatically—always.
- Observability granularity vs. overhead. Persisting full inputs and outputs for every tool call makes storage and retrieval costs nontrivial on long trajectories. The usual compromise: full span metadata for everything + sampled retention for large payloads.
- Leaderboard comparability vs. business relevance. Public benchmarks guarantee comparability, but your business distribution is worlds apart from SWE-bench's. Use public leaderboards to pick models and calibrate; a private eval set is what drives day-to-day iteration—neither can replace the other.
Further reading
- Token consumption in traces feeds directly into the optimization loop of Context Engineering
- Observability instrumentation for tool calls is covered in Tool System; step counts and loop detection echo Agent Loop
- To build a minimal harness with evals yourself, see Build Your Own Harness; eval-related anti-patterns are collected in Common Pitfalls
- For an index of the original papers behind this page's methodology, see Core Papers
- On the case-study side: the different evaluation philosophies of Claude Code and SWE-agent are covered in Claude Code case study and SWE-agent case study
References
- SWE-bench official leaderboard and variant docs (Verified 500 tasks, Lite 300 tasks, etc.)
- SWE-bench Pro paper (arXiv 2509.16941)—public-set SOTA around 44%
- Terminal-Bench site and leaderboard; Terminal-Bench 2.0 notes (Snorkel mirror)—89 tasks, Docker-containerized, designed at launch to hold frontier models below 50%; llm-stats Terminal-Bench 2.0 leaderboard (2026-08: top ~82.7%, average ~58%)
- LangChain blog: lifting Terminal-Bench 2.0 from 52.8 to 66.5 through harness engineering
- τ-bench (arXiv 2406.12045) and τ²-bench paper (arXiv 2506.07982)—the pass^k metric, dual-control environments, GPT-4.1 at 34% on telecom; τ-bench official leaderboard
- WebArena paper (arXiv 2307.13854)—812 tasks, GPT-4 agent 14.41% vs. humans 78.24%; WebArena project site
- WorkArena paper (arXiv 2403.07718)—33 tasks/19,912 instances, GPT-4o 42.7%; BrowserGym paper
- Holistic Agent Leaderboard (arXiv 2510.11977)—the same model across different frameworks spans up to 7 points on GAIA
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (arXiv 2306.05685)—systematic measurement of position bias, verbosity bias, and self-preference
- OpenTelemetry GenAI semantic conventions; Langfuse; LangSmith