Skip to content

Agent Evals from Scratch

At a glance The complete hands-on path to building agent evals: what the unit, trajectory, and outcome layers each assert; how eval datasets grow out of real failures and why 20 cases are enough to start; where programmatic assertions end and LLM-as-judge begins; how to wire evals into CI for regression while trading off cost against sampling—and finally, a minimal eval script for mini_harness.

Agent Evals from Scratch ​

Evaluation & Observability answered the "what": why agent evals are hard, how trajectories get recorded, and what the mainstream benchmarks each measure. This page answers the "how": suppose you've just finished writing a harness (say, the one from Build a Minimal Harness from Scratch)—starting from zero, how do you put together your first eval suite?

The most immediate reason to take this seriously comes from the job market: in this site's survey of agent job descriptions at home and abroad (see JD Checklist), "evaluation / evals" is the most frequently cited hard-skill keyword—Cursor has a dedicated "Agent Evaluation and Quality" role, Tencent Hunyuan is hiring an "Agent Eval Infra Engineer," and Anthropic's FDE job listing puts "build evals that actually capture what matters" first among its responsibilities. Plenty of people can write agents; few can measure them.

Where this page stands

Evaluation is not an acceptance step bolted on after the harness is built; it is the other half of the codebase, growing alongside the harness. A harness change with no eval suite behind it is, in essence, Pitfall #8: Demo-driven development—declaring victory on the strength of one good demo. Every step on this page can be put in place in a single afternoon, with no platform accounts required.

Why it's hard: the one-page version ​

The full argument lives in Evaluation & Observability; here it is compressed into three conclusions with direct consequences for the hands-on work:

  1. Non-determinism → you can't run it just once. The same task against the same harness can produce different results on two runs. A single pass/fail carries no information; your metrics must be distributions (mean, variance, pass^k). This directly drives the engineering requirement below that every case run at least N times.
  2. Trajectory vs. outcome → you can't test just one layer. A correct final result doesn't mean the process was healthy (you may have gotten lucky), and a wrong result doesn't mean everything is lost (maybe only the last step picked the wrong parameter). Look only at the end state and you can't tell where the harness needs fixing.
  3. Long tail → you can't test exhaustively. Traditional testing aims to cover code paths with test cases; an agent's input space is natural language × environment state, which is effectively infinite. The goal of an eval set shifts from "coverage" to "sampling the real distribution"—and that determines how the dataset should be built (expanded on below).

The three layers of evals: what each one asserts ​

Break the vague ask of "evaluate the agent" into three layers; what each layer asserts, how you write it, and what it costs are all different:

text
┌────────────────────────────────────────────────────────────────┐
│  Outcome    Did the task succeed? (end state: files/tests/DB)  │
│    ▲        Objective, fully automatable, but can't say "why"  │
│  Trajectory Is the process healthy? (steps, efficiency, no     │
│    ▲        loops, no hallucinated calls). Pinpoints harness   │
│             defects; fits programmatic heuristic rules         │
│  Unit       Was a single tool call right? (tool name, args,    │
│    ▲        timing). Cheapest, most deterministic — the        │
│             "unit test" of the agent world                     │
└────────────────────────────────────────────────────────────────┘
LayerAsserts onTypical checksTraditional-testing analogueWhen to write
UnitOne model decision / tool callGiven the context, the model should call read("config.yaml") rather than bash("cat ..."); args match the schema; no hallucinated tool namesUnit testWhen the tool protocol changes
TrajectoryThe step sequence of one full runRead before write; no 3 identical calls in a row; steps/cost within budget; no referencing files that were never readIntegration test + linterWhen process-level bad cases show up
OutcomeThe task's end statePatch passes the tests; output file exists with correct content; database state matches expectationsEnd-to-end acceptanceRequired for every eval case

Two practical points:

  • Outcome-level is the backbone; the other two layers are diagnostic tools. Every eval case needs an outcome-level check, otherwise it doesn't count as a test; unit- and trajectory-level assertions are added as needed, and their main value is telling you which layer a failure came from.
  • Unit-level assertions are the most easily overlooked, yet the best value for the effort. A large share of low-level agent failures—malformed args, hallucinated tool names, using bash cat where read was called for—can be caught at the single-call level, without running a full trajectory. Writing assertions on tool calls is like adding compiler checks to the harness's protocol layer.

Step 1: Build the eval dataset ​

Recycle real failures; don't invent questions in a vacuum ​

This is the highest-ROI discipline in the entire evaluation system: for every production bad case, your first instinct should be to freeze it into an eval case. A user reports "the agent corrupted the config file"—save the task description, the environment snapshot, and the failed trajectory from that moment, pair them with a grader asserting "the config file should have been modified correctly," and you have a new case. The distribution of real failures is your business's distribution; test questions made up off the top of your head are not.

The recycling mechanism can be minimal: add a tag: failed marker to the harness's trajectory logs (JSONL), walk through the flagged trajectories once a week, and promote the ones worth keeping into eval cases. The key is not the process; it's muscle memory—a bad case that never makes it into the eval set is a failure wasted.

Three traps of synthetic data ​

Generating eval questions in bulk with an LLM is tempting, but there are three empirically documented traps:

  • Distorted difficulty. Model-generated questions skew toward the "looks like it tests ability but is actually solvable in one step" shape, with far less discriminating power than real tasks. A 90% success rate measured on synthetic questions might be only 60% on the real distribution.
  • Distribution shift. Synthetic data reflects the generating model's "imagination distribution," not your users' distribution. Use GPT to write questions for your own product and you are measuring "what GPT finds hard."
  • Unreliable answers. The labeled answers on synthetic questions may themselves be wrong (the generating model's hallucinations flow straight into the ground truth), so the grader will judge correct behavior as failure—this kind of "bad question" is worse than no question at all, because it rewards wrong behavior.

WARNING

Synthetic data can be used to expand an eval set (perturbing around real cases, generating variants); it must not be used to cold-start one. The first batch of cases has to come from real trajectories or hand-crafted realistic tasks.

Size: start with 20 cases ​

Don't wait until you've "stockpiled 500 cases." Twenty carefully chosen real cases—covering the 3-5 task types you run most often, each type including known failures and known successes—are enough to start exposing problems in the harness. The reason is plain: early harness iterations don't face a statistical-significance problem; they face a "how did nobody catch such an obvious bug" problem. Running 20 cases on every change is worth far more than 500 cases that never get run.

Let the pace of growth follow your iteration stage: 20 cases to cold-start → 50-100 for daily regression → several hundred plus repeated sampling for release gating. The holdout discipline (keeping a batch of questions you never inspect in detail, to guard against overfitting) was covered in Evaluation & Observability and applies here too.

Step 2: Write the graders ​

Programmatic assertions vs. LLM-as-judge: where to draw the line ​

Programmatic assertionsLLM-as-judge
Applies toObjective end states: tests pass, file exists, command output matches, schema validOpen-ended output: summary quality, answer relevance, code readability
CostExpensive to write, free to run, zero varianceQuick to write, billed per token to run, has variance
BiasNone (but has coverage blind spots)Position bias, verbosity bias, self-preference (well documented in the literature; see Evaluation & Observability)
RoleThe sole basis for release gatingRelative comparison and rough screening; conclusions need human spot-checks

The one-line principle: whenever a judgment can be programmatic, make it programmatic. "report.md exists and contains the right number of lines" does not need an LLM to decide; grep is enough. Only pay the cost and variance of LLM-as-judge when the judgment itself requires semantic understanding.

Even when you do use an LLM judge, follow the discipline from Evaluation & Observability: binary judgments item by item on a rubric, pairwise preferred over absolute scoring, judge and judgee on models from different vendors, and 5-10% human spot-checks.

Grader first — and test the grader itself ​

The right order for adding an eval case is counterintuitive: write the grader first, then fix the harness. Concretely—when a new case enters the suite, first confirm that the grader marks the known failure as a failure and the known success as a success (validate by replaying historical trajectories); only then is the grader "calibrated." An uncalibrated grader mixed into the eval set means you are no longer testing the agent but the grader's bugs. This is the same idea as "write the failing test first" in traditional testing.

Step 3: Add an eval script to mini_harness ​

Enough theory; time to build. The .scratch/mini_harness.py from Build a Minimal Harness from Scratch has a clean structure: agent_loop(llm, task) accepts any llm callable (a MockLLM or a wrapper around a real API), and every tool call goes through run_tool. These two seams are the mounting points for the eval script: swap the llm to inject the model under test, and wrap run_tool to record trajectories.

python
# eval_mini.py — a minimal eval script for mini_harness (run from the same directory as mini_harness.py)
import json, os
import mini_harness as mh

# Eval set: task + environment setup + programmatic grader (one example per assertion layer)
CASES = [
    {
        "id": "wc-report",
        "task": "Count the lines in notes.txt and write a report.md",
        "setup": lambda: open("notes.txt", "w").write("a\nb\nc\n"),
        # Outcome level: artifact exists and its content is correct
        "grade_outcome": lambda t: os.path.exists("report.md")
                                   and "3" in open("report.md").read(),
        # Trajectory level: the file was read first, and no tool errored
        "grade_trace": lambda t: any(c["tool"] == "read" for c in t)
                                 and not any("error" in c["result"] for c in t),
        # Unit level: no bash cat standing in for read (protocol-discipline assertion)
        "grade_unit": lambda t: not any(
            c["tool"] == "bash" and "cat " in c["args"].get("command", "") for c in t),
    },
    # ... real cases flow back in from failed trajectories; build up to 20 to start
]

N_RUNS = 3  # Non-determinism: run each case multiple times; look at pass rate, not a single run

def run_once(case, llm):
    """Run once: wrap run_tool to record the trajectory, restore when done."""
    trace, orig = [], mh.run_tool
    def recording(name, args):
        result = orig(name, args)
        trace.append({"tool": name, "args": args, "result": result})
        return result
    mh.run_tool = recording
    try:
        case["setup"]()
        mh.agent_loop(llm, case["task"])   # llm is injected by the caller: real API or MockLLM
    finally:
        mh.run_tool = orig
    return trace

def main(make_llm):
    os.environ["AUTO_APPROVE"] = "1"       # Unattended in CI; skip the approval gate
    for case in CASES:
        passes = 0
        for _ in range(N_RUNS):
            t = run_once(case, make_llm())
            ok = (case["grade_outcome"](t) and case["grade_trace"](t)
                  and case["grade_unit"](t))
            passes += ok
        print(f"{case['id']}: {passes}/{N_RUNS} passed")
        # Key cases should also log their trajectory to JSONL for failure attribution and dataset feedback

if __name__ == "__main__":
    main(lambda: mh.MockLLM([  # Use the MockLLM first to validate the eval script itself; swap in a real model here
        {"tool": "read", "args": {"path": "notes.txt"}},
        {"tool": "bash", "args": {"command": "wc -l notes.txt"}},
        {"tool": "write", "args": {"path": "report.md", "content": "3 lines total\n"}},
        {"done": "done"},
    ]))

This script deliberately keeps three "meta-structures of eval engineering" that hold no matter which framework you switch to:

  • Graders are bound to cases. Each case carries its own setup plus the three layers of grade_*; adding a case is adding one dict—the eval set can therefore be reviewed, diffed, and versioned like code.
  • Test the eval script with a MockLLM first. Before wiring in a real model, get things running end to end with a scripted MockLLM: this confirms the graders themselves work (they pass a "known-good trajectory"). This step is the minimal implementation of "grader calibration" from the previous section.
  • N_RUNS is not optional. Once a real model is wired in, raise it to 3-5 and report the k/N pass rate; a green single run means nothing.

Step 4: Wire evals into CI for regression ​

Once the eval set has accumulated, the biggest waste is running it "only when you remember to." Two kinds of changes must trigger evals automatically:

  • Harness changes: every commit that touches the system prompt, tool descriptions, context-assembly logic, or permission rules. This is exactly the half you actually control, per Model vs. Harness—every time you touch it, have scores to back it up.
  • Model upgrades: switching model versions, switching vendors, even silent updates from the same vendor (API backend drift is real). Run the same eval set before and after the upgrade; the delta is the net effect of the upgrade—the only way you are entitled to believe "the new model is better."

Run in tiers to keep CI duration in check (the tiering idea comes from the trade-off discussion in Evaluation & Observability):

TierSizeTriggerOn failure
smoke10-20 core cases, single sampleEvery commitBlock the merge
regressionFull set, 3-5 samplesMerge to main / model upgradeBlock the release; attribute by hand
fullFull set + holdout, more samplesMilestonesPublish a report; update the baseline

Only three engineering details matter for CI integration: API keys go through secrets management; eval artifacts (trajectory JSONL + summary report) are archived as build artifacts so that a failure lets you open the trajectory and attribute it directly; and the pass-rate threshold is set to no worse than the current baseline, not some absolute number picked out of the air—the point of a gate is preventing regressions, not pursuing perfection.

The cost/sampling trade-off ​

Evals burn money, so run the numbers first: 100 cases × 5 samples each × $0.50 per run on average (a typical magnitude for a moderately complex coding task) = $250 per round. Ten rounds a day is unacceptable. Four levers for cutting cost, in recommended order:

  1. Cut samples before cutting cases. Cases cover the distribution; sample counts only affect confidence—dropping half your cases hurts more than going from 5 samples down to 3.
  2. Tiered triggers (table above): the expensive full evals run only on merges and milestones; everyday commits run only smoke.
  3. Trajectory-level heuristic rules are nearly free. Programmatic checks such as loop detection, context utilization, and cost-over-budget run on every trajectory at zero cost, yet catch most common ailments—let them run on everything, and sample only 10-20% of trajectories for the expensive LLM judge.
  4. Institutionalize judge spot-checking. Have humans review 5-10% of LLM-as-judge conclusions; this is both cost control and bias control.

TIP

Budget eval cost as an inherent overhead of the harness instead of cutting it once the bill spirals—what gets cut is usually exactly the part you need most (long-tail cases and repeated sampling). An honest baseline: eval spending is roughly 10-20% of what you spend running agents in production.

Tooling ​

The script above demonstrates the mechanics; you don't have to rewrite all of it yourself. Three proven options, ordered by how much they take over:

ToolFormBest for
promptfooOpen-source CLI (MIT); cases + assertions defined in YAML; native CI supportThe first step up from this page's script: move CASES into YAML and get assertions, matrix comparisons (multiple models × multiple prompts), and CI gating out of the box
inspect_aiOpen-source Python framework (MIT), built by the UK AI Security InstituteSerious multi-step agent evaluation: Task = dataset + solver + scorer, scorers ranging from exact match to model-graded, with sandboxed execution and full trajectory logs
BraintrustCommercial eval platformTeam-scale workflows: one-click conversion of production trajectories into datasets, experiment diffing, trace-level scoring—"bad case recycling" as a product feature

Only one piece of selection advice: get the mechanics working first, then reach for tools. Run the full loop of "dataset + graders + CI trigger" with this page's 60-line script, and you'll know exactly which part of a tool you actually need—what most teams lack was never a platform; it's the first 20 cases grown out of real failures.

Further reading ​

References ​

  • promptfoo (GitHub) and CI/CD integration docs — open-source, MIT-licensed, CLI-first eval and red-teaming tool
  • inspect_ai (GitHub) — the UK AI Security Institute's Task/Solver/Scorer evaluation framework, a common foundation for system-card evals at frontier labs
  • Braintrust docs — the commercial platform for turning production trajectories into datasets and comparing experiments
  • OpenAI Evals (GitHub) — an early open-source eval framework and registry, worth consulting for how it organizes cases and graders