Appearance
Agent Evals from Scratch
Evaluation & Observability answered the "what": why agent evals are hard, how trajectories get recorded, and what the mainstream benchmarks each measure. This page answers the "how": suppose you've just finished writing a harness (say, the one from Build a Minimal Harness from Scratch)—starting from zero, how do you put together your first eval suite?
The most immediate reason to take this seriously comes from the job market: in this site's survey of agent job descriptions at home and abroad (see JD Checklist), "evaluation / evals" is the most frequently cited hard-skill keyword—Cursor has a dedicated "Agent Evaluation and Quality" role, Tencent Hunyuan is hiring an "Agent Eval Infra Engineer," and Anthropic's FDE job listing puts "build evals that actually capture what matters" first among its responsibilities. Plenty of people can write agents; few can measure them.
Where this page stands
Evaluation is not an acceptance step bolted on after the harness is built; it is the other half of the codebase, growing alongside the harness. A harness change with no eval suite behind it is, in essence, Pitfall #8: Demo-driven development—declaring victory on the strength of one good demo. Every step on this page can be put in place in a single afternoon, with no platform accounts required.
Why it's hard: the one-page version
The full argument lives in Evaluation & Observability; here it is compressed into three conclusions with direct consequences for the hands-on work:
- Non-determinism → you can't run it just once. The same task against the same harness can produce different results on two runs. A single pass/fail carries no information; your metrics must be distributions (mean, variance, pass^k). This directly drives the engineering requirement below that every case run at least N times.
- Trajectory vs. outcome → you can't test just one layer. A correct final result doesn't mean the process was healthy (you may have gotten lucky), and a wrong result doesn't mean everything is lost (maybe only the last step picked the wrong parameter). Look only at the end state and you can't tell where the harness needs fixing.
- Long tail → you can't test exhaustively. Traditional testing aims to cover code paths with test cases; an agent's input space is natural language × environment state, which is effectively infinite. The goal of an eval set shifts from "coverage" to "sampling the real distribution"—and that determines how the dataset should be built (expanded on below).
The three layers of evals: what each one asserts
Break the vague ask of "evaluate the agent" into three layers; what each layer asserts, how you write it, and what it costs are all different:
text
┌────────────────────────────────────────────────────────────────┐
│ Outcome Did the task succeed? (end state: files/tests/DB) │
│ ▲ Objective, fully automatable, but can't say "why" │
│ Trajectory Is the process healthy? (steps, efficiency, no │
│ ▲ loops, no hallucinated calls). Pinpoints harness │
│ defects; fits programmatic heuristic rules │
│ Unit Was a single tool call right? (tool name, args, │
│ ▲ timing). Cheapest, most deterministic — the │
│ "unit test" of the agent world │
└────────────────────────────────────────────────────────────────┘| Layer | Asserts on | Typical checks | Traditional-testing analogue | When to write |
|---|---|---|---|---|
| Unit | One model decision / tool call | Given the context, the model should call read("config.yaml") rather than bash("cat ..."); args match the schema; no hallucinated tool names | Unit test | When the tool protocol changes |
| Trajectory | The step sequence of one full run | Read before write; no 3 identical calls in a row; steps/cost within budget; no referencing files that were never read | Integration test + linter | When process-level bad cases show up |
| Outcome | The task's end state | Patch passes the tests; output file exists with correct content; database state matches expectations | End-to-end acceptance | Required for every eval case |
Two practical points:
- Outcome-level is the backbone; the other two layers are diagnostic tools. Every eval case needs an outcome-level check, otherwise it doesn't count as a test; unit- and trajectory-level assertions are added as needed, and their main value is telling you which layer a failure came from.
- Unit-level assertions are the most easily overlooked, yet the best value for the effort. A large share of low-level agent failures—malformed args, hallucinated tool names, using
bash catwherereadwas called for—can be caught at the single-call level, without running a full trajectory. Writing assertions on tool calls is like adding compiler checks to the harness's protocol layer.
Step 1: Build the eval dataset
Recycle real failures; don't invent questions in a vacuum
This is the highest-ROI discipline in the entire evaluation system: for every production bad case, your first instinct should be to freeze it into an eval case. A user reports "the agent corrupted the config file"—save the task description, the environment snapshot, and the failed trajectory from that moment, pair them with a grader asserting "the config file should have been modified correctly," and you have a new case. The distribution of real failures is your business's distribution; test questions made up off the top of your head are not.
The recycling mechanism can be minimal: add a tag: failed marker to the harness's trajectory logs (JSONL), walk through the flagged trajectories once a week, and promote the ones worth keeping into eval cases. The key is not the process; it's muscle memory—a bad case that never makes it into the eval set is a failure wasted.
Three traps of synthetic data
Generating eval questions in bulk with an LLM is tempting, but there are three empirically documented traps:
- Distorted difficulty. Model-generated questions skew toward the "looks like it tests ability but is actually solvable in one step" shape, with far less discriminating power than real tasks. A 90% success rate measured on synthetic questions might be only 60% on the real distribution.
- Distribution shift. Synthetic data reflects the generating model's "imagination distribution," not your users' distribution. Use GPT to write questions for your own product and you are measuring "what GPT finds hard."
- Unreliable answers. The labeled answers on synthetic questions may themselves be wrong (the generating model's hallucinations flow straight into the ground truth), so the grader will judge correct behavior as failure—this kind of "bad question" is worse than no question at all, because it rewards wrong behavior.
WARNING
Synthetic data can be used to expand an eval set (perturbing around real cases, generating variants); it must not be used to cold-start one. The first batch of cases has to come from real trajectories or hand-crafted realistic tasks.
Size: start with 20 cases
Don't wait until you've "stockpiled 500 cases." Twenty carefully chosen real cases—covering the 3-5 task types you run most often, each type including known failures and known successes—are enough to start exposing problems in the harness. The reason is plain: early harness iterations don't face a statistical-significance problem; they face a "how did nobody catch such an obvious bug" problem. Running 20 cases on every change is worth far more than 500 cases that never get run.
Let the pace of growth follow your iteration stage: 20 cases to cold-start → 50-100 for daily regression → several hundred plus repeated sampling for release gating. The holdout discipline (keeping a batch of questions you never inspect in detail, to guard against overfitting) was covered in Evaluation & Observability and applies here too.
Step 2: Write the graders
Programmatic assertions vs. LLM-as-judge: where to draw the line
| Programmatic assertions | LLM-as-judge | |
|---|---|---|
| Applies to | Objective end states: tests pass, file exists, command output matches, schema valid | Open-ended output: summary quality, answer relevance, code readability |
| Cost | Expensive to write, free to run, zero variance | Quick to write, billed per token to run, has variance |
| Bias | None (but has coverage blind spots) | Position bias, verbosity bias, self-preference (well documented in the literature; see Evaluation & Observability) |
| Role | The sole basis for release gating | Relative comparison and rough screening; conclusions need human spot-checks |
The one-line principle: whenever a judgment can be programmatic, make it programmatic. "report.md exists and contains the right number of lines" does not need an LLM to decide; grep is enough. Only pay the cost and variance of LLM-as-judge when the judgment itself requires semantic understanding.
Even when you do use an LLM judge, follow the discipline from Evaluation & Observability: binary judgments item by item on a rubric, pairwise preferred over absolute scoring, judge and judgee on models from different vendors, and 5-10% human spot-checks.
Grader first — and test the grader itself
The right order for adding an eval case is counterintuitive: write the grader first, then fix the harness. Concretely—when a new case enters the suite, first confirm that the grader marks the known failure as a failure and the known success as a success (validate by replaying historical trajectories); only then is the grader "calibrated." An uncalibrated grader mixed into the eval set means you are no longer testing the agent but the grader's bugs. This is the same idea as "write the failing test first" in traditional testing.
Step 3: Add an eval script to mini_harness
Enough theory; time to build. The .scratch/mini_harness.py from Build a Minimal Harness from Scratch has a clean structure: agent_loop(llm, task) accepts any llm callable (a MockLLM or a wrapper around a real API), and every tool call goes through run_tool. These two seams are the mounting points for the eval script: swap the llm to inject the model under test, and wrap run_tool to record trajectories.
python
# eval_mini.py — a minimal eval script for mini_harness (run from the same directory as mini_harness.py)
import json, os
import mini_harness as mh
# Eval set: task + environment setup + programmatic grader (one example per assertion layer)
CASES = [
{
"id": "wc-report",
"task": "Count the lines in notes.txt and write a report.md",
"setup": lambda: open("notes.txt", "w").write("a\nb\nc\n"),
# Outcome level: artifact exists and its content is correct
"grade_outcome": lambda t: os.path.exists("report.md")
and "3" in open("report.md").read(),
# Trajectory level: the file was read first, and no tool errored
"grade_trace": lambda t: any(c["tool"] == "read" for c in t)
and not any("error" in c["result"] for c in t),
# Unit level: no bash cat standing in for read (protocol-discipline assertion)
"grade_unit": lambda t: not any(
c["tool"] == "bash" and "cat " in c["args"].get("command", "") for c in t),
},
# ... real cases flow back in from failed trajectories; build up to 20 to start
]
N_RUNS = 3 # Non-determinism: run each case multiple times; look at pass rate, not a single run
def run_once(case, llm):
"""Run once: wrap run_tool to record the trajectory, restore when done."""
trace, orig = [], mh.run_tool
def recording(name, args):
result = orig(name, args)
trace.append({"tool": name, "args": args, "result": result})
return result
mh.run_tool = recording
try:
case["setup"]()
mh.agent_loop(llm, case["task"]) # llm is injected by the caller: real API or MockLLM
finally:
mh.run_tool = orig
return trace
def main(make_llm):
os.environ["AUTO_APPROVE"] = "1" # Unattended in CI; skip the approval gate
for case in CASES:
passes = 0
for _ in range(N_RUNS):
t = run_once(case, make_llm())
ok = (case["grade_outcome"](t) and case["grade_trace"](t)
and case["grade_unit"](t))
passes += ok
print(f"{case['id']}: {passes}/{N_RUNS} passed")
# Key cases should also log their trajectory to JSONL for failure attribution and dataset feedback
if __name__ == "__main__":
main(lambda: mh.MockLLM([ # Use the MockLLM first to validate the eval script itself; swap in a real model here
{"tool": "read", "args": {"path": "notes.txt"}},
{"tool": "bash", "args": {"command": "wc -l notes.txt"}},
{"tool": "write", "args": {"path": "report.md", "content": "3 lines total\n"}},
{"done": "done"},
]))This script deliberately keeps three "meta-structures of eval engineering" that hold no matter which framework you switch to:
- Graders are bound to cases. Each case carries its own
setupplus the three layers ofgrade_*; adding a case is adding one dict—the eval set can therefore be reviewed, diffed, and versioned like code. - Test the eval script with a MockLLM first. Before wiring in a real model, get things running end to end with a scripted MockLLM: this confirms the graders themselves work (they pass a "known-good trajectory"). This step is the minimal implementation of "grader calibration" from the previous section.
N_RUNSis not optional. Once a real model is wired in, raise it to 3-5 and report thek/Npass rate; a green single run means nothing.
Step 4: Wire evals into CI for regression
Once the eval set has accumulated, the biggest waste is running it "only when you remember to." Two kinds of changes must trigger evals automatically:
- Harness changes: every commit that touches the system prompt, tool descriptions, context-assembly logic, or permission rules. This is exactly the half you actually control, per Model vs. Harness—every time you touch it, have scores to back it up.
- Model upgrades: switching model versions, switching vendors, even silent updates from the same vendor (API backend drift is real). Run the same eval set before and after the upgrade; the delta is the net effect of the upgrade—the only way you are entitled to believe "the new model is better."
Run in tiers to keep CI duration in check (the tiering idea comes from the trade-off discussion in Evaluation & Observability):
| Tier | Size | Trigger | On failure |
|---|---|---|---|
| smoke | 10-20 core cases, single sample | Every commit | Block the merge |
| regression | Full set, 3-5 samples | Merge to main / model upgrade | Block the release; attribute by hand |
| full | Full set + holdout, more samples | Milestones | Publish a report; update the baseline |
Only three engineering details matter for CI integration: API keys go through secrets management; eval artifacts (trajectory JSONL + summary report) are archived as build artifacts so that a failure lets you open the trajectory and attribute it directly; and the pass-rate threshold is set to no worse than the current baseline, not some absolute number picked out of the air—the point of a gate is preventing regressions, not pursuing perfection.
The cost/sampling trade-off
Evals burn money, so run the numbers first: 100 cases × 5 samples each × $0.50 per run on average (a typical magnitude for a moderately complex coding task) = $250 per round. Ten rounds a day is unacceptable. Four levers for cutting cost, in recommended order:
- Cut samples before cutting cases. Cases cover the distribution; sample counts only affect confidence—dropping half your cases hurts more than going from 5 samples down to 3.
- Tiered triggers (table above): the expensive full evals run only on merges and milestones; everyday commits run only smoke.
- Trajectory-level heuristic rules are nearly free. Programmatic checks such as loop detection, context utilization, and cost-over-budget run on every trajectory at zero cost, yet catch most common ailments—let them run on everything, and sample only 10-20% of trajectories for the expensive LLM judge.
- Institutionalize judge spot-checking. Have humans review 5-10% of LLM-as-judge conclusions; this is both cost control and bias control.
TIP
Budget eval cost as an inherent overhead of the harness instead of cutting it once the bill spirals—what gets cut is usually exactly the part you need most (long-tail cases and repeated sampling). An honest baseline: eval spending is roughly 10-20% of what you spend running agents in production.
Tooling
The script above demonstrates the mechanics; you don't have to rewrite all of it yourself. Three proven options, ordered by how much they take over:
| Tool | Form | Best for |
|---|---|---|
| promptfoo | Open-source CLI (MIT); cases + assertions defined in YAML; native CI support | The first step up from this page's script: move CASES into YAML and get assertions, matrix comparisons (multiple models × multiple prompts), and CI gating out of the box |
| inspect_ai | Open-source Python framework (MIT), built by the UK AI Security Institute | Serious multi-step agent evaluation: Task = dataset + solver + scorer, scorers ranging from exact match to model-graded, with sandboxed execution and full trajectory logs |
| Braintrust | Commercial eval platform | Team-scale workflows: one-click conversion of production trajectories into datasets, experiment diffing, trace-level scoring—"bad case recycling" as a product feature |
Only one piece of selection advice: get the mechanics working first, then reach for tools. Run the full loop of "dataset + graders + CI trigger" with this page's 60-line script, and you'll know exactly which part of a tool you actually need—what most teams lack was never a platform; it's the first 20 cases grown out of real failures.
Further reading
- Evaluation & Observability — the conceptual foundation for this page: why it's hard, the trajectory schema, a tour of benchmarks, LLM-as-judge bias
- Build a Minimal Harness from Scratch — home of this page's eval script, and the same "components must earn their place" philosophy
- Pitfalls & Antipatterns — demo-driven development is the very disease evals are meant to cure
- Harness Design Principles — which changes eval gates should attach to
- JD Checklist — the raw evidence behind evals as the most frequent job requirement
References
- promptfoo (GitHub) and CI/CD integration docs — open-source, MIT-licensed, CLI-first eval and red-teaming tool
- inspect_ai (GitHub) — the UK AI Security Institute's Task/Solver/Scorer evaluation framework, a common foundation for system-card evals at frontier labs
- Braintrust docs — the commercial platform for turning production trajectories into datasets and comparing experiments
- OpenAI Evals (GitHub) — an early open-source eval framework and registry, worth consulting for how it organizes cases and graders