Skip to content

Common Pitfalls & Anti-Patterns

At a glance Ten high-frequency anti-patterns in agent harness development — framework-first design, tool sprawl, context hoarding, no termination conditions, swallowed errors, full-auto everything, prompt silver bullets, demo-driven development, blind spots on cost and latency, and premature multi-agent — each with symptoms, root cause, and the fix.

Common Pitfalls & Anti-Patterns ​

Harnesses fail very differently from traditional software. When traditional software breaks, it throws an exception, prints a stack trace, and lights up a dashboard. When a harness breaks, it usually just gets quietly dumber — no errors, just output that keeps getting worse, loops that keep getting longer, and bills that keep getting bigger. By the time you hear about it from user complaints, the problem has usually been there for weeks.

This shapes what harness anti-patterns look like: most of them aren't "written wrong," they're "missing one constraint." This page collects the ten most common pitfalls from community practice as a list, each broken into Symptoms → Root cause → The fix, with example scenarios that stay as close to real life as possible.

How to use this page

Skim the quick reference table below first; when a symptom matches, jump to the corresponding section. If you're building a harness from scratch, you can also use this page as a review checklist.

Quick Reference Table ​

#Anti-patternTypical symptomOne-line fix
1Framework-firstChanging one behavior means reading three layers of abstractionGet it working with a bare API + a while loop first
2Tool sprawlAgent picks the wrong tool and fills in garbage parametersConsolidate into a small set of orthogonal tools; load the rest on demand
3Context hoardingThe longer the session, the dumber and more expensive it getsTruncate/compress/offload every tool output
4No termination conditionsRunaway loop burns tokens; the bill lands at 3 a.m.Hard caps on steps, cost, and time
5Swallowed errorsAgent repeats the same failed operation ten timesErrors must flow back into the context, structured
6Full-auto everythingAgent drops the production database; nobody signed offGate operations by how irreversible they are
7Prompt silver bullet3,000-line system prompt and it's still misbehavingWhen the fix is structural, fix the structure
8Demo-driven development"Look, I'll show you it works"Write evals before talking about capabilities
9Blind to cost & latency$8 per task, 40 seconds to first responseTreat cost as a first-class metric in observability
10Premature multi-agentFive agents passing the buck to each otherPush single-agent + subtask isolation to its limits first

1. Framework-First ​

Symptoms

The project introduces a heavyweight agent framework on day one, and two weeks later the team's main job has become: reading the framework's source code, filing issues against the framework, and working around the framework's default behavior. You want to tweak one retry policy and find it buried three callbacks deep; you want to switch models and discover the framework's abstractions block provider-specific parameters.

Root cause

Projecting an "illusion of maturity" onto the framework: assuming that what it encapsulates is stable domain knowledge. In reality, best practices in the agent space have been rewritten roughly every six months since 2023, and what a framework encapsulates is usually last quarter's consensus. Anthropic's observation after working with dozens of customer teams is blunt: the most successful implementations use simple, composable patterns rather than complex frameworks (see Building effective agents, 2024-12).

Example scenario

Team A built a code-review agent on some framework's PlanAndExecuteAgent class. After launch they found the planner's plans were always over-decomposed, and they wanted to change it to "plan only two steps at a time." But the planner was a built-in framework component and the prompt template was a private attribute, so in the end the only option was to monkey-patch half the class — by which point the framework patches they maintained were longer than their actual business code.

The fix

  • The first version of your harness gets exactly three things: an LLM API call, a while loop, and a tool dispatch table. A few dozen lines of code and it runs. See Build Your Own Harness.
  • You can adopt a framework — but only when it solves a problem you have already felt the pain of. For example, only consider a stateful-graph framework like LangGraph once you genuinely need production-grade checkpoint persistence (see the LangGraph case study).
  • The test: can you draw, on a single sheet of paper, what the framework does for you? If you can't, the framework is making decisions for you.

One lesson

The framework APIs you learn become worthless the moment you switch frameworks. The agent loop, context assembly, and error handling you learn in a bare loop carry over to any project. Master those first.


2. Tool Sprawl ​

Symptoms

The agent has 50 tools: read_file, read_lines, read_file_range, read_file_with_encoding... and then it starts doing things: calling the wrong tool, filling tool A's parameters into tool B, bouncing back and forth between three overlapping tools, or forcing a tool call when none is needed.

Root cause

Every tool you add makes the model's selection problem one notch harder, and every tool's schema occupies the context window. The research evidence here is remarkably consistent: more candidate tools actively degrade selection accuracy (How Many Tools Should an LLM Agent See?, arXiv 2605.24660). The RAG-MCP paper names "stuffing every tool into the prompt" prompt bloat outright and mitigates it with retrieval-based selection (arXiv 2505.03275). The QLCoder team's engineering retrospective records the same thing: handing the agent a large set of heterogeneous tools caused selection chaos, and behavior stabilized immediately once they converged on a small, well-defined toolbox (arXiv 2511.08462).

Example scenario

An ops agent was wired up to 40 internal APIs. While investigating a "service 502," it first called get_service_status (came back healthy), then list_pods (parameters filled in wrong), then get_metrics (picked the wrong time window), and after twenty minutes delivered its conclusion: "all normal." After the toolkit was trimmed to 8 tools and three status queries were merged into one inspect_service call with a scope parameter, similar tasks converged in 6 steps on average.

The fix

  • Merge overlapping tools: every "read a file" tool becomes one tool, with parameters expressing the differences. For sound tool design principles, see Tool Design.
  • Mount tools dynamically per task phase: read-only tools during planning; write tools only during execution. The toolset is per-step, not per-agent.
  • When you genuinely need a large tool library, add a retrieval/routing layer (RAG-MCP style) so the model sees 5–10 candidates per step instead of the entire registry.
  • Follow SWE-agent's lead: its ACI (agent-computer interface) deliberately narrows the interface down to a small set of high-cohesion commands, and the paper specifically demonstrates that interface design alone significantly affects success rates (SWE-agent paper, NeurIPS 2024; see also the SWE-agent case study).

3. Context Hoarding ​

Symptoms

By turn 30 of a session, the agent starts "losing memory": it forgets constraints set at the start, re-reads files it has already read, and the quality of its conclusions visibly slides. Open the trace and you find an 8,000-line npm install output, three full SQL dumps, and the entire history of some git log sitting in the context.

Root cause

Two mental errors stack up: "the window is big enough, it'll fit," and "what if we need it later?" But context is not a hard drive — every token spends the model's attention budget. Liu et al.'s Lost in the Middle (arXiv 2307.03172, 2023) showed that models' use of mid-context information collapses in a U shape; Anthropic later gave the phenomenon an engineering name, context rot: as the window fills up, recall accuracy keeps degrading (Effective context engineering for AI agents, 2025-09).

Example scenario

[step 12] $ grep -rn "TODO" src/          → returns 2,300 lines, all of it into context
[step 13] $ cat package-lock.json          → 41,000 lines, all of it into context
[step 14] $ python manage.py test          → 600-line failure traceback, into context
...
[step 30] model: I see the problem is... (has long forgotten the user's step 1 instruction: "don't touch the db directory")

The fix

  • Every tool output passes through a gate: truncate (head/tail), summarize, or write to a temp file and put only the path and a summary back into the context. This is the core move of Context Engineering.
  • Separate "references" from "content": keep lightweight pointers in the context — file paths, line ranges, query parameters — and let the model fetch what it needs when it needs it. Just-in-time retrieval, not loading everything up front.
  • Compact long sessions: summarize history into "established facts + open items" and continue in a fresh window. This is exactly the strategy Claude Code uses; see the Claude Code case study.

A counterintuitive point

"Make the agent remember everything" is not a virtue. A human engineer debugging closes the irrelevant terminal tabs; a harness should do the same for the agent.


4. Runaway Loop ​

Symptoms

Some Friday night, the agent picks up an edge-case task and spins in "try → fail → rephrase and try again → fail" for four hours, burning $300. Saturday morning you stare at the usage dashboard wondering whether the system has come under attack.

Root cause

The agent loop is fundamentally while not done: step(), and done is judged by the model itself. Models are unreliable at self-assessing on hard tasks — it always feels like "one more try and we're there." A ReAct-style loop (ReAct paper, ICLR 2023) carries no built-in notion of budget; the budget has to be forced in at the harness layer.

The fix

Termination conditions must be multiple layers of hard constraints, hard-coded into the loop, relying on nothing from the model's own judgment:

python
MAX_STEPS = 50
MAX_COST_USD = 2.0
MAX_WALL_CLOCK = 1800  # seconds

while True:
    if state.steps >= MAX_STEPS:
        return fail("step budget exhausted", partial=state)
    if state.cost_usd >= MAX_COST_USD:
        return fail("cost budget exhausted", partial=state)
    if time.monotonic() - state.started_at > MAX_WALL_CLOCK:
        return fail("time budget exhausted", partial=state)
    step(state)

Add one behavior-level check on top: N consecutive steps with no state change (same tool, same arguments, or an empty diff) is declared a dead loop and exits immediately. When the budget runs out, the return value should carry the partial state — partial progress is usually worth more than "start over." For more on termination design, see The Agent Loop.


5. Swallowed Errors ​

Symptoms

The logs say the task "completed successfully," but the output is wrong. Replay the trace and you find it: the file write in step 4 actually failed, and the API call in step 7 returned 403 — but the agent knew none of this and went on stacking twenty more blocks on a broken premise.

Root cause

Two classic ways to write it:

python
# Anti-pattern A: swallow the exception
try:
    result = run_tool(name, args)
except Exception:
    result = None          # what the model sees is "no output", not "it failed"

# Anti-pattern B: downgrade the error to vague text
except Exception as e:
    result = "something went wrong"   # the model cannot correct its behavior based on this

If the model never receives the "failure" signal, it won't retry or change course. Worse, it will treat the failed step's empty output as "the operation succeeded but produced nothing" and build that into its subsequent reasoning.

The fix

  • Inject errors back into the context, structured: error type, error message, a (sanitized) stack summary, a suggested direction for the fix. Make "failure" a first-class input the model can reason about.
  • Distinguish retryable errors (network timeouts, rate limits) from non-retryable errors (permission denied, invalid parameters) — the harness auto-retries the former with backoff; for the latter it must tell the model plainly: "this path is a dead end, try something else."
  • Give tool returns a uniform envelope: { ok: bool, output?: string, error?: { kind, message, hint } }. Anthropic's SWE-bench practice notes the same thing: they spent more time polishing their tools — including how errors get reported back — than tuning prompts (Building effective agents).
What a good error injection looks like
json
{
  "ok": false,
  "error": {
    "kind": "permission_denied",
    "message": "EACCES: cannot write to /etc/nginx/nginx.conf",
    "hint": "This path requires root. Options: 1) write to ./nginx.conf and ask the user to deploy it manually; 2) if sudo is available, set sudo=true when calling run_shell"
  }
}

Handed this, the model can make a meaningful correction on the very next step. Handed null, all it can do is guess.


6. Full-Auto Everything ​

Symptoms

To demo "fully unattended end-to-end," in its first week live the agent: git push --force overwrote a colleague's commits, rm -rf'd a directory that wasn't under version control yet, and replied to a public issue with the contents of an internal discussion. Every one of those traces shows the model "doing what seemed reasonable at the time."

Root cause

Treating autonomy as a boolean instead of a dial. In reality, operations differ in danger across at least three tiers: read-only < reversible writes < irreversible/outbound. Applying one approval policy to all three tiers means staking the highest-risk tier on the model's judgment at every single step.

The fix

Gate by irreversibility and tighten tier by tier:

TierExamplesDefault policy
L0 read-onlyRead files, search, query statusAuto-approve
L1 reversible writesEdit workspace files, run testsAuto-approve, with an audit log
L2 hard to reverseDelete files, modify the database, commit codeHuman confirmation by default; configurable whitelist
L3 outbound / moneySend email, comment on PRs, call paid APIs, deploy to productionAlways require human confirmation

This is exactly the subject the Permissions, Safety & Human-in-the-Loop chapter covers. Two more engineering moves that are cheap and pay off big: a dry-run mode (the agent first shows a human "here's what I plan to do") and an impact preview (before deleting, list the files that would be affected).

TIP

"Requires human confirmation" is not distrust of the model's ability — it is respect for the distribution of error costs: the price of missing one bad action is far greater than the price of ninety-nine extra confirmations.


7. Prompt-Only Tuning ​

Symptoms

The system prompt has bloated from 200 words to 3,000, stuffed with "Very important:" "Never:" "Let me say this again:" — but the agent hasn't gotten better, and new, stranger failures have appeared. Every bad case adds one more prohibition to the prompt, until the prompt becomes a rulebook nobody (including the model) can finish reading.

Root cause

Many problems have no solution at the prompt layer at all:

"agent keeps picking the wrong tool"     → root cause is a 50-tool choice space, not the wording
"agent forgets early constraints"        → root cause is context rot, not too few reminders
"agent mangles parameter formats"        → root cause is ambiguous tool schema design, not too few examples
"agent runs in circles"                  → root cause is no termination condition; shouting in the prompt won't help

Patching the prompt is, in essence, converting structural debt into context noise — and accelerating context rot along the way.

The fix

Build a disciplined troubleshooting order:

  1. Read the trace first; locate which layer the failure lives in: the toolset? context assembly? termination conditions? permissions? observability?
  2. Whatever can be fixed at the structural layer should not be fixed at the prompt layer. Too many tools: cut tools. Dirty context: add a gate. Runaway loop: add a budget.
  3. The prompt handles only what cannot be expressed structurally: domain preferences, tone, the soft edges of judgment calls.
  4. For every prompt rule you add, ask: "Six months from now, after a model swap, will this rule still be needed?" If the answer is mostly "no," it probably belonged in the harness all along.

The SWE-agent paper offers a positive reference point here: the key to their success-rate gains was not a better prompt but a redesign of the interface between the agent and the computer (the ACI) — making correct behavior easier and incorrect behavior harder (arXiv 2405.15793).


8. Demo-Driven Development ​

Symptoms

Capability is proven by "look, let me demo it" — three demos, two successes, screenshot the successful one. After launch, when users report bad cases, the team's fix process is: reproduce manually → tweak something → reproduce manually again → pray nothing new broke. No number anywhere can answer "is this week's version better or worse than last week's?"

Root cause

Agent behavior is probabilistic; a single demonstration carries almost no information. Without evals, every change is groping in the dark: prompt tweaks, model upgrades, adding or removing tools — each one might improve one cluster of tasks while collapsing another, and you have no instrument to see it happen.

The fix

  • Build an eval set on day one, even if it's only 20 cases: real tasks + expected results + a checker that can score automatically (did the tests pass? does the diff match? another model as judge?).
  • Hook evals into CI: every prompt change, model swap, or tool change triggers a run; watch three curves — success rate, average steps, average cost.
  • The destination for bad cases is the eval set, not the chat log — turn every production failure into a regression case.
  • This and Evaluation & Observability are two sides of one coin: traces show you a single run; evals show you the overall trend.

A pragmatic starting point

Don't wait for the "perfect eval framework." A JSONL file plus a script that prints a pass rate at the end already beats where most teams stand. Get numbers first; talk about the quality of the numbers after.


9. Blind to Cost & Latency ​

Symptoms

Acceptance testing is all green — until finance or a user comes knocking: $8 average per task, 40-second P50 first response; the support agent takes three planning steps, two tool calls, and 90 seconds to answer "what are your opening hours?"

Root cause

During development, only one metric was watched: success rate. But agent systems burn money multiplicatively with step count: every step re-sends (or re-summarizes) the whole context, so token consumption grows roughly as the square of the number of turns, and every step's latency lands directly on the user's wait time. One figure worth remembering from Anthropic's context engineering article: agents consume tokens at a scale far beyond ordinary chat, and multi-agent architectures amplify that further (Effective context engineering for AI agents).

The fix

  • Put cost into the trace: every span records input/output tokens, model, unit price, and duration; roll up per task into the three metrics "cost / steps / duration," on the same dashboard as success rate.
  • Tiered routing: route simple requests to a small model, or even deterministic code; save the big model for genuinely hard problems. Routing is one of the foundational patterns in Building effective agents and the best bang-for-buck optimization on this list.
  • Cache everything you can: the system prompt and tool schemas are stable — use prompt caching; cache repeated retrieval results by key.
  • Parallelize independent tool calls; don't serialize steps that could run in parallel.
  • Set budget alerts (echoing #4): when a single task's cost crosses the threshold, prefer degrading to "partial result + explanation" over grinding all the way to the end.

10. Premature Multi-Agent ​

Symptoms

The single agent isn't even tuned yet, but the architecture diagram already shows five roles: Orchestrator, Researcher, Coder, Reviewer, Critic. At runtime, the five agents each see only their own context: the Coder doesn't know what the Researcher found, the Reviewer is criticizing code built on a wrong premise, and the Orchestrator dispatches the same task twice. The token bill is several times a single agent's, and the success rate is lower.

Root cause

Multi-agent splits the context, not the capability. Sub-agents don't share working memory, so all coordination happens through explicit message passing — and every handoff is a lossy compression. Cognition (the Devin team) makes the point sharply in Don't Build Multi-Agents (2025-06, Walden Yan): when parallel sub-agents each make their own decisions, the moment their implicit assumptions diverge, the merge stage produces conflicts that cannot be reconciled; their principle is "share context, share the full trace."

Example scenario

A team split their data-analysis agent into a "data-fetching agent" and a "charting agent." The data-fetching agent found one field was all nulls and switched to a different table — a decision that existed only in its trace. The charting agent then carefully laid out a chart against "the wrong table" and wrote an interpretation to go with it. To locate the divergence, a human had to replay both traces side by side; debugging cost doubled.

The fix

  • First push single-agent + good context management to its ceiling: compaction, just-in-time retrieval, structured notes (see Memory Systems). A lot of "we need multiple agents" is really "our single agent's context is dirty."
  • Genuine signals that multi-agent applies: the subtasks are parallelizable, loosely coupled, and independently verifiable (breadth-first research is the classic case). If those three don't hold, splitting only amplifies the coordination tax.
  • The halfway option is the Subagents pattern: a sub-agent does focused work in a clean window and returns only a compressed conclusion to the main agent — you keep the benefit of context isolation while decision-making stays in one place.
  • Read the counterargument before deciding: LangChain's How and when to build multi-agent systems (2025-06) responds to Cognition's post and draws the boundary of applicability fairly on both sides.

Mapping Anti-Patterns to Harness Components ​

When troubleshooting, use this table to work backwards to "which chapter should I read":

Anti-patternPrimary component
Framework-firstOverall architecture — see Anatomy of the Harness
Tool sprawl, swallowed errors, prompt-only tuning (partly)Tool Design
Context hoarding, prompt-only tuning (partly)Context Engineering
No termination conditionsThe Agent Loop
Full-auto everythingPermissions, Safety & Human-in-the-Loop
Demo-driven development, blind to cost & latencyEvaluation & Observability
Premature multi-agentSubagents

Three Universal Defense Principles ​

Abstract all the anti-patterns up a level and there are only three root diseases — and only three defenses:

  1. Everything goes into the trace, and the trace goes into evals. Silence is the harness's number-one enemy. A run you can't see cannot be debugged, and an improvement you can't quantify is not an improvement.
  2. Write constraints into the structure, not into the prayers. Budgets, permissions, termination conditions, output gates — anything "the model shouldn't do" should be something the model cannot do, not something it was told not to do.
  3. Complexity is debt; decide how much to borrow by how fast the interest accrues. Frameworks, multi-agent, long prompts, big toolboxes — every one of them is borrowed complexity. Before borrowing, ask: is the interest it charges (debugging cost, context noise, coordination tax) below the value it delivers?

Further Reading ​

  • Design Principles: the positive version of this page — anti-patterns tell you what not to do; design principles tell you what to do
  • Build Your Own Harness: step through the shallow versions of these pitfalls yourself with a minimal implementation
  • Model vs. Harness: why most of these pitfalls have little to do with model capability
  • Classic Papers, Annotated: full reading guides for the papers cited here — ReAct, Lost in the Middle, SWE-agent, and more

References ​