Appearance
Building an Agent Eval Suite from Scratch
Evaluation Systems laid out the theory and the layered methodology of Agent evals; the 30-line evals.py at the end of the Progressive Tutorial gave you a minimal baseline. This page is the middle ground between the two: take the tutorial's V2 research assistant (v2_researcher.py) as the system under test and build, from zero, an eval suite that can live in CI.
When you're done you'll have:
- A 20-50 task set with a difficulty gradient and contamination safeguards;
- A layered grader combining code assertions with LLM-as-judge (with a complete, usable judge prompt);
- Automatic per-run logging of trajectories, cost, and latency, generating a Markdown report;
- A GitHub Actions regression workflow: run before and after prompt changes, and block anything that loses points.
Methodologically, this page leans primarily on Anthropic's January 2026 engineering post Demystifying Evals for AI Agents—the most systematic public write-up on Agent evals from the Claude Code team. Several conclusions here (the capability/regression split, pass@k/pass^k, grader design principles) come directly from it and are noted where used.
1. Setting the Scene: the System Under Test and Shared Vocabulary
The system under test is the tutorial's research assistant: receive a topic → update_todo splits sub-questions → loop over web_search → read_url → save_note → read_all_notes to consolidate → write_report produces workspace/report.md. Six tools, max_turns=25.
First, align on terminology (these words come from Anthropic's eval guide and are broadly industry-standard):
- task: one test case = input + success criteria;
- trial: one execution of a task. Model output is stochastic, so a task runs multiple trials;
- grader: scoring logic. A task can carry multiple graders, each containing multiple assertions (checks);
- transcript: the full record of one trial (every message, every tool call and its result), also called a trace / trajectory;
- outcome: the final state of the environment when the trial ends. For the research assistant, the outcome is the actual content of
workspace/report.md, not the Agent's closing line "report generated."
Capability evals and regression evals are two different things
The most important distinction in Anthropic's guide: capability evals ask "what can this Agent do well"—pick tasks it's bad at; low pass rates are normal; they're a mountain to climb. Regression evals ask "can it still do what it could do before"—pass rates should be near 100%; any drop means something broke. What we build here is mainly a regression suite (20-50 previously "passing" tasks guarding against degradation), plus a small set of capability tasks (deliberately hard ones) to drive improvement. When a capability task's pass rate rises and stabilizes, "graduate" it into the regression suite.
The whole pipeline looks like this; each section below implements one stage:
tasks.jsonl ──► eval_runner ──► run k trials per task
│ (isolated workspace/)
▼
┌── transcript logging: messages, tool calls, tokens, timing
│
▼
layered grader ──► Layer 1: code assertions (structure/keywords/tool sequence)
│ failure → straight FAIL, no judge money spent
▼
Layer 2: LLM-as-judge (coverage/groundedness/source quality)
│
▼
results.jsonl + report.md ──► compare against baseline ──► CI gate2. Task Set Design: What 20-50 Tasks Are Really About
The task set is the asset most worth your time in the entire eval system. You can write the script in two days; the task set takes months to cultivate.
The difficulty gradient
Split research-assistant tasks into three tiers by "what capability they demand," each with different failure modes:
| Tier | Capability demanded | Example | Suggested share |
|---|---|---|---|
| L1 single fact | one search, one concrete fact answered correctly | "What was company X's Q3 2025 revenue?" | ~30% |
| L2 multi-source comparison | search multiple sources, structured comparison | "Compare LangGraph vs the OpenAI Agents SDK for selection scenarios" | ~50% |
| L3 open survey | multi-round research, synthesis, surfacing controversies | "Summarize the mainstream 2026 Agent eval methods and compare their strengths" | ~20% |
L1's expected output is close to a closed answer and tolerates strong assertions; L3 can only be given a "must-cover points list" and leans on the judge. The distribution deliberately sits on L2—that's the research assistant's home turf.
What a task looks like
Store the task set as JSONL, one task per line:
json
{"id": "l2-frameworks-01", "level": "L2",
"input": "Compare the applicable scenarios of LangGraph and the OpenAI Agents SDK and give a selection recommendation",
"must_cover": ["LangGraph orchestrates via explicit state graphs", "The OpenAI Agents SDK has fewer abstractions and is faster to pick up",
"Choose LangGraph for complex orchestration", "Choose the Agents SDK for rapid prototyping"],
"min_sources": 3, "expect_tools": ["web_search", "read_url", "save_note", "write_report"],
"max_turns": 25}Design rules—every one of them earned the hard way or sourced:
- Every task must pass on its own first. Anthropic's words: a task should be one that "an agent that correctly follows instructions will definitely pass." Run each new task by hand once and confirm your mental "perfect answer" passes all graders—that's your reference solution, and it also validates that the graders weren't miswritten.
must_coverlists "points," not "sentences." The judge checks points; the Agent can phrase them any way it likes. Graders that hard-code sentences misfire on legitimate variants (on CORE-Bench, Opus 4.5 got crushed to 42 because the grader rejected "96.12" for "96.124991…"; with the grader fixed it scored 95—overly rigid graders can cut a model's real capability in half).- Balance tasks in both directions. Include "should search" tasks and also "shouldn't search—should refuse or say information is insufficient from common knowledge" tasks (e.g., "research a company that doesn't exist"). A one-directional task set trains a one-directional Agent. When Anthropic built search evals for Claude.ai, they deliberately covered both directions: things to search (weather) and things not to (who founded Apple).
- Difficulty has to be high enough. A task set that's all passes carries no information. Keep a few L3 stumpers the current Agent can't pass—that's the capability-eval portion.
Where tasks come from, and contamination control
Where do tasks come from? By priority:
- Real usage logs: cases where the research assistant failed on you (turn them into tasks immediately);
- Hand-written: authored against the Agent's capability boundary;
- Model-generated + human curation: have an LLM draft tasks in bulk, then pick—most efficient, but a human must review every one; model-drafted tasks skew easy and formulaic.
Contamination has three layers:
- Don't let tasks live where the Agent can read them. The research assistant searches the web; if your task set sits in a public repo, it can theoretically find the "answers." Keep the task-set repo private, or at least store the
must_coverpoints separately from the question text. - Don't write tasks into prompts. Putting "for example, you can compare LangGraph and the Agents SDK" into the instructions is leaking the exam.
- Research ground truths go stale. Anthropic's guide specifically notes that research Agents' ground truths drift with web content. "Company X's latest revenue" will be wrong three months later. Countermeasures: prefer structural points over concrete numbers in
must_cover; tag pure-fact questions with anexpiresdate and refresh or retire them on schedule.
Tasks: only add, never modify—expand regularly
If pass rate rises after a prompt change but your gut says nothing improved, suspect task-set overfitting—you unconsciously tuned the prompts to "just barely pass the existing tasks." Discipline: only add tasks, never rewrite existing ones; add 3-5 new tasks after each major Agent iteration; graduate capability tasks into the regression suite once they pass consistently. The tutorial mentioned this trap too; it's repeated here because it is genuinely the #1 cause of eval systems going rotten.
3. Grader Implementation: Code Assertions + LLM-as-judge, Layered
The principle in one sentence (Anthropic's words): wherever a deterministic grader can do the job, don't use an LLM judge. Code assertions are fast, free, reproducible, and easy to debug; judges are expensive, stochastic, and need calibrating against human judgment. So: layer them. Run code assertions first; any failure is an immediate FAIL, and the judge only sees trials that clear the code layer.
Layer 1: code assertions
For a research assistant, determinism can check more than you'd expect:
python
# graders_code.py
import re
from pathlib import Path
def grade_structure(workspace: Path, task: dict) -> list[str]:
"""Structural assertions; returns a list of failure reasons (empty = all pass)."""
failures = []
report_path = workspace / "report.md"
if not report_path.exists():
return ["report.md not generated"]
report = report_path.read_text(encoding="utf-8")
# 1. The report isn't an empty shell: minimum length
if len(report) < 500:
failures.append(f"report too short ({len(report)} chars)")
# 2. Source citations: count the http(s) links in the text
urls = re.findall(r"https?://[^\s)\]]+", report)
if len(urls) < task["min_sources"]:
failures.append(f"insufficient sources: {len(urls)} < {task['min_sources']}")
# 3. Required sections (structural points, not exact wording)
for section in ["Summary", "References"]:
if section not in report:
failures.append(f"missing '{section}' section")
# 4. Plan and notes actually landed on disk (process-artifact assertions)
if not (workspace / "plan.md").exists():
failures.append("plan.md not generated (planning skipped)")
notes = list((workspace / "notes").glob("*.md")) if (workspace / "notes").exists() else []
if not notes:
failures.append("no notes saved (research process skipped)")
return failuresNote check 4: it queries the real state of the environment (do the files exist), not what the Agent said in the transcript. That's outcome grading—Anthropic's example: a flight-booking Agent saying "booked!" doesn't count; the row in the SQL database does.
Tool-call sequence assertions live at the trajectory layer—see section 4.
Layer 2: the LLM-as-judge prompt template
Code assertions can't check "is the report actually right, is it complete." That's the judge's turf. Three design principles (all from Anthropic's guide and judge best practices):
- Score each dimension in its own call—never all dimensions in one prompt—cram five dimensions into one prompt and the judge's scores interfere with each other;
- Give the judge an "insufficient information" exit, otherwise it's forced to bluff a judgment;
- Require reasons before scores (critique-then-score), with the score last in the JSON—judges that emit a score first will invent reasons around it.
Here's the complete judge prompt for the coverage dimension, ready to use:
text
You are a strict reviewer of research reports. Evaluate how well the report below covers the required points.
# Research topic
{task_input}
# Required-coverage point list
{must_cover_bullets}
# Full research report
{report}
# Review instructions
1. For each point, judge whether the report covers it: "covered" (explicitly discussed), "partial" (mentioned but insufficient),
or "missing" (not addressed).
2. Judge only from the report text, not from your own knowledge. If the report's content is insufficient to judge, mark "unknown".
3. First output the per-point analysis (evidence quotes the report verbatim), then the aggregate score.
4. Aggregate score: coverage_score = covered count / total points (partial counts 0.5).
# Output format (strict JSON, no other content)
{
"checks": [
{"point": "...", "verdict": "covered|partial|missing|unknown", "evidence": "..."}
],
"coverage_score": 0.0
}Write the other two dimensions with the same structure:
- Groundedness: give the judge the report plus the list of source URLs collected in the transcript; ask it to spot-check 5 concrete factual claims from the report and rule on whether each has source support;
- Source quality: require a judgment on whether the cited sources are "authoritative primary sources (official docs/papers/filings)" or "secondary reposts"; deduct when the authoritative share falls below a threshold.
The judge itself needs evaluating
LLM-as-judge isn't trustworthy just because you finished writing it. Anthropic recommends calibrating against human expert judgment: sample 20-30 already-graded items for human review and measure judge–human agreement; if agreement is too low, fix the prompt (usually the dimension definitions aren't concrete enough). Hamel Husain's complete LLM-as-a-Judge guide covers this calibration workflow in unusual depth—worth a read. One more engineering trick: run the judge at temperature=0, and the judge model doesn't have to be the strongest—a cheap model is often enough for point-checking, and the savings buy more trials.
How the two layers combine
A trial's final score:
- Any code assertion fails →
score = 0, judge not called (saves money, and a high judge score on top of a code-layer failure is meaningless); - Code layer fully passes → the judge scores the three dimensions, combined by weight:
score = 0.4 * coverage + 0.4 * groundedness + 0.2 * source_quality; - Task-level ruling:
score >= 0.7counts as PASS. This is Anthropic's weighted scoring; if your team is more conservative, use binary (all graders must pass) instead—pick one and stick with it, don't flip back and forth.
4. Trajectory Analysis: How the Agent Got to Its Result
Scoring outcomes alone misses a whole class of problems: the result happens to be right, but the process is wrong (detours, repeated calls, a lucky hit on the final turn). Trajectory analysis examines the process. For the research assistant: four metrics and one assertion:
| Metric / assertion | Healthy signal | Danger signal |
|---|---|---|
| Number of steps (turns) | 8-20 turns | hitting the max_turns=25 cap and getting truncated |
| Tool-call sequence | a repeating web_search → read_url → save_note cycle | never calling read_url at all (sources never read) |
| Repeated calls | occasional | same URL read 3 times in a row, same query searched repeatedly |
| Tool failure rate | <10% of calls return errors | lots of "fetch failed" with no strategy adjustment afterward |
"Stuck detection" needs just one simple rule: if the transcript shows 3 consecutive calls with the same tool + same arguments, mark stuck=True. More reliable than any LLM judgment, and free.
Tool sequence assertions (extending the code-assertion layer from section 3):
python
def grade_trajectory(tool_calls: list[str], task: dict) -> list[str]:
"""tool_calls is the list of tool names in chronological order."""
failures = []
for t in task["expect_tools"]: # key tools must have appeared
if t not in tool_calls:
failures.append(f"never called {t}")
# stuck detection: 3 identical calls in a row
for i in range(len(tool_calls) - 2):
if tool_calls[i] == tool_calls[i + 1] == tool_calls[i + 2]:
failures.append(f"possibly stuck: {tool_calls[i]} called 3 times in a row")
break
return failuresNote that the "expected tools" assertion checks only appearance, not order or count. An Agent finding a legitimate path you didn't foresee is the norm—grade what it produced, not which road it took. This is Anthropic's "Don't check for specific paths" put into practice.
Where does the trajectory data come from: the OpenAI Agents SDK's RunResult.new_items contains every item from the run in order; entries with type == "tool_call_item" are tool calls (tool name at item.raw_item.name); token usage is at result.context_wrapper.usage. Field names may shift between SDK versions—defer to the official docs; the script in section 7 uses getattr defensively. Heavier trajectory needs (visualization, connecting production traces with evals) belong on platforms like LangSmith / Langfuse—see Observability.
5. Cost and Latency Logging
Have the eval script log each trial's cost and latency while it's at it—that's your decision data for future cost optimization. Record at trial granularity; when aggregating, look at distributions, not just means:
- Per trial:
n_turns,n_toolcalls,input_tokens,output_tokens, wall-clock time; - Per eval run: total tokens, total cost converted via the price table, P50/P95 latency;
- Trend: append one row per run to
results.jsonl, so cost spikes jump out at a glance.
Don't hard-code the price table in the script—model prices change. Keep a pricing.json (model name → per-million input/output token prices) and copy current values from each vendor's pricing page. Both extended uses of evals come from this record: did the pass rate drop after switching to a cheaper model (evals tell you whether it was worth it); average turns rose from 12 to 20 after some change (you probably broke the prompt, even if the pass rate held).
On non-determinism, two metrics you must know (core content of Anthropic's guide):
- pass@k: the probability of at least one pass in k attempts. Rises with k; measures "one lucky hit is enough" scenarios;
- pass^k: the probability that all k attempts pass. Falls with k; measures "stable every time" scenarios—user-facing Agents should be judged on this. An Agent with a 75% single-run pass rate has a pass^3 of only about 42%.
Engineering implication: run 3 trials per task and report both pass@1 (average single-run pass rate) and pass^3 (share of tasks that passed all three). When comparing before/after a change, look at pass^3—it's far more sensitive to degradation.
6. CI Integration: Let Evals Block Bad Changes
Evals outside CI are just a toy. Workflow design:
- Trigger: a PR touching
v2_researcher.py, theinstructions, or tool code triggers the eval workflow; full runs are too expensive, so gate with a PR label (only run with therun-evalslabel). - Gate: if the regression suite's pass rate drops more than a threshold (say 10 percentage points) versus the
mainbranch baseline, the workflow exits non-zero and the PR turns red. - Report: generate a Markdown summary, posted as a PR comment or CI artifact—the author immediately sees which tasks dropped and what failure reasons the judge gave.
yaml
# .github/workflows/agent-evals.yml
name: agent-evals
on:
pull_request:
paths: ["v2_researcher.py", "evals/**"]
jobs:
evals:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with: { python-version: "3.12" }
- run: pip install -U openai-agents httpx
- name: Download baseline
uses: actions/download-artifact@v4
with: { name: eval-baseline, path: evals/baseline/ }
continue-on-error: true # no baseline on the first run
- name: Run evals (3 trials × full task set)
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
run: python evals/eval_suite.py --trials 3 --baseline evals/baseline/results.jsonl
- name: Upload report and new baseline
uses: actions/upload-artifact@v4
with: { name: eval-report, path: evals/out/ }eval_suite.py's job: run the evals → compare against baseline → exit 1 when degradation exceeds the threshold. The report looks like this (evals/out/report.md, paste straight into the PR):
text
## Eval Report · run 2026-08-21T09:30 · trials=3
| Metric | baseline | this run | Δ |
| --- | --- | --- | --- |
| pass@1 | 82% | 74% | -8pp |
| pass^3 | 71% | 58% | -13pp ❌ |
| avg turns | 11.8 | 16.2 | +4.4 |
| total cost | $1.42 | $2.10 | +48% |
Regressed tasks:
- l2-frameworks-01: coverage 0.83→0.50 (judge: missing the "choose LangGraph for complex orchestration" point)
- l3-eval-methods-02: code assertion failed "insufficient sources: 2 < 3"The comparison discipline: change exactly one variable per commit (prompt, or tools, or model), run the evals, record. Change three things at once and a rise can't be attributed and a fall can't be blamed—the same mistake as tuning without a baseline.
The pragmatic trade-offs of eval frequency
A full eval run (30 tasks × 3 trials × multi-turn tool calls) can take 20-40 minutes and a few dollars of API spend—not for every commit. Common practice: run a smoke subset of 10 core tasks on PRs; run the full set nightly on a schedule; run the full set plus a human transcript spot-check before releases. This matches the "offline evals + online monitoring, layered" thinking in Evaluation Systems.
7. The Complete Eval Script
Assembling all the parts above. Single file; the only dependencies are openai-agents and httpx (the tested Agent's own deps); the judge goes through an OpenAI-compatible client directly:
python
# evals/eval_suite.py
"""Research assistant eval script: task set → multi-trial runs → layered grading → report and gate."""
import argparse
import asyncio
import json
import re
import shutil
import sys
import tempfile
import time
from pathlib import Path
from agents import Runner
from openai import OpenAI
from v2_researcher import agent # system under test
from graders_code import grade_structure # code assertions from section 3
JUDGE_MODEL = "gpt-5.6-luna" # judge model; good enough beats strongest
JUDGE_PROMPT = (Path(__file__).parent / "judge_coverage.md").read_text(encoding="utf-8")
client = OpenAI()
def extract_trajectory(result) -> tuple[list[str], dict]:
"""Extract the tool-call sequence and usage from a RunResult (defensive access, in case SDK fields shift)."""
tools = [getattr(i.raw_item, "name", "?") for i in result.new_items
if getattr(i, "type", "") == "tool_call_item"]
usage = getattr(result.context_wrapper, "usage", None)
metrics = {
"n_turns": len(result.new_items),
"n_toolcalls": len(tools),
"input_tokens": getattr(usage, "input_tokens", 0),
"output_tokens": getattr(usage, "output_tokens", 0),
}
return tools, metrics
def grade_trajectory(tool_calls: list[str], task: dict) -> list[str]:
failures = []
for t in task.get("expect_tools", []):
if t not in tool_calls:
failures.append(f"never called {t}")
for i in range(len(tool_calls) - 2):
if tool_calls[i] == tool_calls[i + 1] == tool_calls[i + 2]:
failures.append(f"possibly stuck: {tool_calls[i]} 3 times in a row")
break
return failures
def judge_coverage(task: dict, report: str) -> float:
"""LLM-as-judge: coverage score, 0-1. Returns 0 conservatively on failure."""
prompt = (JUDGE_PROMPT
.replace("{task_input}", task["input"])
.replace("{must_cover_bullets}",
"\n".join(f"- {p}" for p in task["must_cover"]))
.replace("{report}", report))
try:
resp = client.chat.completions.create(
model=JUDGE_MODEL, temperature=0,
response_format={"type": "json_object"},
messages=[{"role": "user", "content": prompt}],
)
return float(json.loads(resp.choices[0].message.content)["coverage_score"])
except Exception as e:
print(f" judge call failed: {e}")
return 0.0
async def run_trial(task: dict, trial_idx: int) -> dict:
"""One trial: isolated workspace → run the Agent → layered grading."""
ws = Path(tempfile.mkdtemp(prefix=f"eval_{task['id']}_{trial_idx}_"))
t0 = time.time()
result = await Runner.run(agent, task["input"], max_turns=task.get("max_turns", 25))
latency = time.time() - t0
# Note: v2_researcher's tools hard-code WORKSPACE=Path("workspace").
# A real implementation should read WORKSPACE_DIR from an environment variable; isolation here uses env vars.
tools, metrics = extract_trajectory(result)
report_path = ws / "report.md"
report = report_path.read_text(encoding="utf-8") if report_path.exists() else ""
failures = grade_structure(ws, task) + grade_trajectory(tools, task)
if failures: # code layer failed: FAIL, judge not called
score = 0.0
else: # code layer fully passed: judge scores (example uses coverage only; extend as needed)
score = judge_coverage(task, report)
record = {"task_id": task["id"], "trial": trial_idx,
"pass": score >= 0.7, "score": round(score, 3),
"failures": failures, "latency_s": round(latency, 1), **metrics}
shutil.rmtree(ws, ignore_errors=True) # burn the isolated environment after use
return record
async def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--tasks", default="evals/tasks.jsonl")
ap.add_argument("--trials", type=int, default=3)
ap.add_argument("--baseline", default=None)
ap.add_argument("--out", default="evals/out")
args = ap.parse_args()
tasks = [json.loads(l) for l in Path(args.tasks).read_text(encoding="utf-8").splitlines() if l.strip()]
out = Path(args.out); out.mkdir(parents=True, exist_ok=True)
records = []
for task in tasks: # run tasks serially so web tools don't interfere with each other
for k in range(args.trials):
r = await run_trial(task, k)
records.append(r)
print(f"[{task['id']} t{k}] {'PASS' if r['pass'] else 'FAIL'} "
f"score={r['score']} turns={r['n_turns']} {r['failures'] or ''}")
# Aggregate: pass@1 = mean pass rate across all trials; pass^k = share of tasks passing every trial
pass1 = sum(r["pass"] for r in records) / len(records)
by_task = {t["id"]: [r for r in records if r["task_id"] == t["id"]] for t in tasks}
passk = sum(all(r["pass"] for r in rs) for rs in by_task.values()) / len(tasks)
total_tokens = sum(r["input_tokens"] + r["output_tokens"] for r in records)
# Persist: results.jsonl (raw records) + report.md (human-readable)
with open(out / "results.jsonl", "a", encoding="utf-8") as f:
for r in records:
f.write(json.dumps(r, ensure_ascii=False) + "\n")
summary = {"pass@1": pass1, f"pass^{args.trials}": passk, "total_tokens": total_tokens}
(out / "report.md").write_text(
f"## Eval Report · trials={args.trials}\n\n"
+ "\n".join(f"- {k}: **{v:.0%}**" if isinstance(v, float) and v <= 1 else f"- {k}: {v}"
for k, v in summary.items()) + "\n\nRegressed tasks:\n"
+ "\n".join(f"- {tid}" for tid, rs in by_task.items()
if any(r["pass"] for r in rs) and not all(r["pass"] for r in rs)),
encoding="utf-8")
print(f"\npass@1={pass1:.0%} pass^{args.trials}={passk:.0%} tokens={total_tokens}")
# Gate: compare against baseline; exit code 1 if pass@1 drops more than 10pp
if args.baseline and Path(args.baseline).exists():
base = [json.loads(l) for l in Path(args.baseline).read_text(encoding="utf-8").splitlines() if l.strip()]
base_pass1 = sum(r["pass"] for r in base) / len(base)
if pass1 < base_pass1 - 0.10:
print(f"❌ Regression: pass@1 {base_pass1:.0%} → {pass1:.0%}")
return 1
return 0
if __name__ == "__main__":
sys.exit(asyncio.run(main()))Two practical notes:
- Workspace isolation is a hard requirement. The script above gives each trial its own temp directory and deletes it afterward. Shared state between Agents creates mirages—Anthropic stepped on this internally: in some eval tasks, Claude "cheated" by reading git history left behind by the previous trial, gaining an unfair advantage. The
WORKSPACEconstant inv2_researcher.pyneeds to read from an environment variable so each trial points at its own temp directory. - Store the judge prompt in a separate file (
judge_coverage.md) and version it like the tested Agent's prompts. Changes to the judge prompt should themselves be re-run against the calibration samples, otherwise you can't tell whether a score change came from the Agent or the judge.
If a task fails in every trial (0% pass@k), suspect the task first
This is the debugging instinct Anthropic's guide keeps repeating: 0% pass@100 almost always means a bug in the task definition or the grader, not a hopeless Agent. Usual culprits: an expected point written wrong, dead URLs leaving read_url empty, min_sources set too harshly. Run the reference solution by hand first; only once it passes do you get to suspect the Agent.
8. An Iteration Story: "Regression → Diagnosis → Fix" Post-mortem
The last section walks through using this system for real (numbers are illustrative):
- The change: someone finds the report too wordy and adds a line to
instructions: "Research should be efficient; read at most 1 source per sub-question." - Evals sound the alarm: in CI, pass@1 falls from 82% to 68%, pass^3 from 71% to 55%; the gate blocks the PR. The report shows the regressions are all L2/L3 tasks, with judge failure reasons clustered on "incomplete point coverage."
- Read transcripts to diagnose: pull one regressed task's trajectory—the Agent obeyed the new instruction strictly, calling
read_urlonly once per sub-question; the problem is the first hit is often a secondary repost missing key data. It used to read 2-3 sources and cross-check; now it can't. - The qualitative conclusion: "read 1 source" wasn't wrong—for L1 single-fact questions it's genuinely faster and cheaper; the mistake was applying it wholesale, restricting tasks that need multi-source cross-validation too.
- The fix: revise the instruction to "read 1 source for single facts; for comparisons, data, or contested claims, read at least 2 independent sources and cross-validate." Re-run the evals: pass@1 84%, with average turns slightly down from before.
- Consolidate: add the newly exposed "trusting a single secondary source" failure mode as a new task in the task set (
l2-crosscheck-01) to prevent future regressions.
The key step in this flow is 3: without transcripts and the judge's per-point failure reasons, all you see is "pass rate dropped," with no idea where to fix. Half an eval system's value is the alarm; the other half is the direction to fix in. That's why Anthropic lists "Read the transcripts" as the final commandment of its guide—scores are only the entry point; transcripts are the facts.
At this point, this eval suite plus the tutorial's research assistant is resume-worthy: "Built a 30-task tiered eval set and a code-assertion + LLM-as-judge two-layer grading pipeline for a research Agent, wired into a CI regression gate; used trajectory analysis to locate and fix a multi-source-verification regression, lifting pass@1 by 16pp"—for how to write it, see Resume Analysis.
References
- Demystifying Evals for AI Agents — Anthropic Engineering — this page's primary methodology source: task/trial/grader terminology, capability vs regression, pass@k/pass^k, grader design principles
- Using LLM-as-a-Judge For Evaluation: A Complete Guide — Hamel Husain — the most detailed public guide to judge prompt writing and human calibration
- LLM Evaluation: Methods, Best Practices, and a Practical Roadmap — Langfuse — survey of LLM-as-judge and eval methods, including a platform-level perspective
- LLM Agent Evaluation Metrics in 2026 — Confident AI — a categorized roundup of tool-call, task-completion, and trajectory metrics
- How to build LLM-as-a-Judge evaluators that hold up in production — Arize — reliability and calibration practices for judges in production
- OpenAI Cookbook: Evaluation examples — OpenAI's official hands-on example of the eval flywheel