Appearance
Common Pitfalls & Anti-Patterns
Harnesses fail very differently from traditional software. When traditional software breaks, it throws an exception, prints a stack trace, and lights up a dashboard. When a harness breaks, it usually just gets quietly dumber — no errors, just output that keeps getting worse, loops that keep getting longer, and bills that keep getting bigger. By the time you hear about it from user complaints, the problem has usually been there for weeks.
This shapes what harness anti-patterns look like: most of them aren't "written wrong," they're "missing one constraint." This page collects the ten most common pitfalls from community practice as a list, each broken into Symptoms → Root cause → The fix, with example scenarios that stay as close to real life as possible.
How to use this page
Skim the quick reference table below first; when a symptom matches, jump to the corresponding section. If you're building a harness from scratch, you can also use this page as a review checklist.
Quick Reference Table
| # | Anti-pattern | Typical symptom | One-line fix |
|---|---|---|---|
| 1 | Framework-first | Changing one behavior means reading three layers of abstraction | Get it working with a bare API + a while loop first |
| 2 | Tool sprawl | Agent picks the wrong tool and fills in garbage parameters | Consolidate into a small set of orthogonal tools; load the rest on demand |
| 3 | Context hoarding | The longer the session, the dumber and more expensive it gets | Truncate/compress/offload every tool output |
| 4 | No termination conditions | Runaway loop burns tokens; the bill lands at 3 a.m. | Hard caps on steps, cost, and time |
| 5 | Swallowed errors | Agent repeats the same failed operation ten times | Errors must flow back into the context, structured |
| 6 | Full-auto everything | Agent drops the production database; nobody signed off | Gate operations by how irreversible they are |
| 7 | Prompt silver bullet | 3,000-line system prompt and it's still misbehaving | When the fix is structural, fix the structure |
| 8 | Demo-driven development | "Look, I'll show you it works" | Write evals before talking about capabilities |
| 9 | Blind to cost & latency | $8 per task, 40 seconds to first response | Treat cost as a first-class metric in observability |
| 10 | Premature multi-agent | Five agents passing the buck to each other | Push single-agent + subtask isolation to its limits first |
1. Framework-First
Symptoms
The project introduces a heavyweight agent framework on day one, and two weeks later the team's main job has become: reading the framework's source code, filing issues against the framework, and working around the framework's default behavior. You want to tweak one retry policy and find it buried three callbacks deep; you want to switch models and discover the framework's abstractions block provider-specific parameters.
Root cause
Projecting an "illusion of maturity" onto the framework: assuming that what it encapsulates is stable domain knowledge. In reality, best practices in the agent space have been rewritten roughly every six months since 2023, and what a framework encapsulates is usually last quarter's consensus. Anthropic's observation after working with dozens of customer teams is blunt: the most successful implementations use simple, composable patterns rather than complex frameworks (see Building effective agents, 2024-12).
Example scenario
Team A built a code-review agent on some framework's
PlanAndExecuteAgentclass. After launch they found the planner's plans were always over-decomposed, and they wanted to change it to "plan only two steps at a time." But the planner was a built-in framework component and the prompt template was a private attribute, so in the end the only option was to monkey-patch half the class — by which point the framework patches they maintained were longer than their actual business code.
The fix
- The first version of your harness gets exactly three things: an LLM API call, a
whileloop, and a tool dispatch table. A few dozen lines of code and it runs. See Build Your Own Harness. - You can adopt a framework — but only when it solves a problem you have already felt the pain of. For example, only consider a stateful-graph framework like LangGraph once you genuinely need production-grade checkpoint persistence (see the LangGraph case study).
- The test: can you draw, on a single sheet of paper, what the framework does for you? If you can't, the framework is making decisions for you.
One lesson
The framework APIs you learn become worthless the moment you switch frameworks. The agent loop, context assembly, and error handling you learn in a bare loop carry over to any project. Master those first.
2. Tool Sprawl
Symptoms
The agent has 50 tools: read_file, read_lines, read_file_range, read_file_with_encoding... and then it starts doing things: calling the wrong tool, filling tool A's parameters into tool B, bouncing back and forth between three overlapping tools, or forcing a tool call when none is needed.
Root cause
Every tool you add makes the model's selection problem one notch harder, and every tool's schema occupies the context window. The research evidence here is remarkably consistent: more candidate tools actively degrade selection accuracy (How Many Tools Should an LLM Agent See?, arXiv 2605.24660). The RAG-MCP paper names "stuffing every tool into the prompt" prompt bloat outright and mitigates it with retrieval-based selection (arXiv 2505.03275). The QLCoder team's engineering retrospective records the same thing: handing the agent a large set of heterogeneous tools caused selection chaos, and behavior stabilized immediately once they converged on a small, well-defined toolbox (arXiv 2511.08462).
Example scenario
An ops agent was wired up to 40 internal APIs. While investigating a "service 502," it first called
get_service_status(came back healthy), thenlist_pods(parameters filled in wrong), thenget_metrics(picked the wrong time window), and after twenty minutes delivered its conclusion: "all normal." After the toolkit was trimmed to 8 tools and three status queries were merged into oneinspect_servicecall with ascopeparameter, similar tasks converged in 6 steps on average.
The fix
- Merge overlapping tools: every "read a file" tool becomes one tool, with parameters expressing the differences. For sound tool design principles, see Tool Design.
- Mount tools dynamically per task phase: read-only tools during planning; write tools only during execution. The toolset is per-step, not per-agent.
- When you genuinely need a large tool library, add a retrieval/routing layer (RAG-MCP style) so the model sees 5–10 candidates per step instead of the entire registry.
- Follow SWE-agent's lead: its ACI (agent-computer interface) deliberately narrows the interface down to a small set of high-cohesion commands, and the paper specifically demonstrates that interface design alone significantly affects success rates (SWE-agent paper, NeurIPS 2024; see also the SWE-agent case study).
3. Context Hoarding
Symptoms
By turn 30 of a session, the agent starts "losing memory": it forgets constraints set at the start, re-reads files it has already read, and the quality of its conclusions visibly slides. Open the trace and you find an 8,000-line npm install output, three full SQL dumps, and the entire history of some git log sitting in the context.
Root cause
Two mental errors stack up: "the window is big enough, it'll fit," and "what if we need it later?" But context is not a hard drive — every token spends the model's attention budget. Liu et al.'s Lost in the Middle (arXiv 2307.03172, 2023) showed that models' use of mid-context information collapses in a U shape; Anthropic later gave the phenomenon an engineering name, context rot: as the window fills up, recall accuracy keeps degrading (Effective context engineering for AI agents, 2025-09).
Example scenario
[step 12] $ grep -rn "TODO" src/ → returns 2,300 lines, all of it into context
[step 13] $ cat package-lock.json → 41,000 lines, all of it into context
[step 14] $ python manage.py test → 600-line failure traceback, into context
...
[step 30] model: I see the problem is... (has long forgotten the user's step 1 instruction: "don't touch the db directory")The fix
- Every tool output passes through a gate: truncate (head/tail), summarize, or write to a temp file and put only the path and a summary back into the context. This is the core move of Context Engineering.
- Separate "references" from "content": keep lightweight pointers in the context — file paths, line ranges, query parameters — and let the model fetch what it needs when it needs it. Just-in-time retrieval, not loading everything up front.
- Compact long sessions: summarize history into "established facts + open items" and continue in a fresh window. This is exactly the strategy Claude Code uses; see the Claude Code case study.
A counterintuitive point
"Make the agent remember everything" is not a virtue. A human engineer debugging closes the irrelevant terminal tabs; a harness should do the same for the agent.
4. Runaway Loop
Symptoms
Some Friday night, the agent picks up an edge-case task and spins in "try → fail → rephrase and try again → fail" for four hours, burning $300. Saturday morning you stare at the usage dashboard wondering whether the system has come under attack.
Root cause
The agent loop is fundamentally while not done: step(), and done is judged by the model itself. Models are unreliable at self-assessing on hard tasks — it always feels like "one more try and we're there." A ReAct-style loop (ReAct paper, ICLR 2023) carries no built-in notion of budget; the budget has to be forced in at the harness layer.
The fix
Termination conditions must be multiple layers of hard constraints, hard-coded into the loop, relying on nothing from the model's own judgment:
python
MAX_STEPS = 50
MAX_COST_USD = 2.0
MAX_WALL_CLOCK = 1800 # seconds
while True:
if state.steps >= MAX_STEPS:
return fail("step budget exhausted", partial=state)
if state.cost_usd >= MAX_COST_USD:
return fail("cost budget exhausted", partial=state)
if time.monotonic() - state.started_at > MAX_WALL_CLOCK:
return fail("time budget exhausted", partial=state)
step(state)Add one behavior-level check on top: N consecutive steps with no state change (same tool, same arguments, or an empty diff) is declared a dead loop and exits immediately. When the budget runs out, the return value should carry the partial state — partial progress is usually worth more than "start over." For more on termination design, see The Agent Loop.
5. Swallowed Errors
Symptoms
The logs say the task "completed successfully," but the output is wrong. Replay the trace and you find it: the file write in step 4 actually failed, and the API call in step 7 returned 403 — but the agent knew none of this and went on stacking twenty more blocks on a broken premise.
Root cause
Two classic ways to write it:
python
# Anti-pattern A: swallow the exception
try:
result = run_tool(name, args)
except Exception:
result = None # what the model sees is "no output", not "it failed"
# Anti-pattern B: downgrade the error to vague text
except Exception as e:
result = "something went wrong" # the model cannot correct its behavior based on thisIf the model never receives the "failure" signal, it won't retry or change course. Worse, it will treat the failed step's empty output as "the operation succeeded but produced nothing" and build that into its subsequent reasoning.
The fix
- Inject errors back into the context, structured: error type, error message, a (sanitized) stack summary, a suggested direction for the fix. Make "failure" a first-class input the model can reason about.
- Distinguish retryable errors (network timeouts, rate limits) from non-retryable errors (permission denied, invalid parameters) — the harness auto-retries the former with backoff; for the latter it must tell the model plainly: "this path is a dead end, try something else."
- Give tool returns a uniform envelope:
{ ok: bool, output?: string, error?: { kind, message, hint } }. Anthropic's SWE-bench practice notes the same thing: they spent more time polishing their tools — including how errors get reported back — than tuning prompts (Building effective agents).
What a good error injection looks like
json
{
"ok": false,
"error": {
"kind": "permission_denied",
"message": "EACCES: cannot write to /etc/nginx/nginx.conf",
"hint": "This path requires root. Options: 1) write to ./nginx.conf and ask the user to deploy it manually; 2) if sudo is available, set sudo=true when calling run_shell"
}
}Handed this, the model can make a meaningful correction on the very next step. Handed null, all it can do is guess.
6. Full-Auto Everything
Symptoms
To demo "fully unattended end-to-end," in its first week live the agent: git push --force overwrote a colleague's commits, rm -rf'd a directory that wasn't under version control yet, and replied to a public issue with the contents of an internal discussion. Every one of those traces shows the model "doing what seemed reasonable at the time."
Root cause
Treating autonomy as a boolean instead of a dial. In reality, operations differ in danger across at least three tiers: read-only < reversible writes < irreversible/outbound. Applying one approval policy to all three tiers means staking the highest-risk tier on the model's judgment at every single step.
The fix
Gate by irreversibility and tighten tier by tier:
| Tier | Examples | Default policy |
|---|---|---|
| L0 read-only | Read files, search, query status | Auto-approve |
| L1 reversible writes | Edit workspace files, run tests | Auto-approve, with an audit log |
| L2 hard to reverse | Delete files, modify the database, commit code | Human confirmation by default; configurable whitelist |
| L3 outbound / money | Send email, comment on PRs, call paid APIs, deploy to production | Always require human confirmation |
This is exactly the subject the Permissions, Safety & Human-in-the-Loop chapter covers. Two more engineering moves that are cheap and pay off big: a dry-run mode (the agent first shows a human "here's what I plan to do") and an impact preview (before deleting, list the files that would be affected).
TIP
"Requires human confirmation" is not distrust of the model's ability — it is respect for the distribution of error costs: the price of missing one bad action is far greater than the price of ninety-nine extra confirmations.
7. Prompt-Only Tuning
Symptoms
The system prompt has bloated from 200 words to 3,000, stuffed with "Very important:" "Never:" "Let me say this again:" — but the agent hasn't gotten better, and new, stranger failures have appeared. Every bad case adds one more prohibition to the prompt, until the prompt becomes a rulebook nobody (including the model) can finish reading.
Root cause
Many problems have no solution at the prompt layer at all:
"agent keeps picking the wrong tool" → root cause is a 50-tool choice space, not the wording
"agent forgets early constraints" → root cause is context rot, not too few reminders
"agent mangles parameter formats" → root cause is ambiguous tool schema design, not too few examples
"agent runs in circles" → root cause is no termination condition; shouting in the prompt won't helpPatching the prompt is, in essence, converting structural debt into context noise — and accelerating context rot along the way.
The fix
Build a disciplined troubleshooting order:
- Read the trace first; locate which layer the failure lives in: the toolset? context assembly? termination conditions? permissions? observability?
- Whatever can be fixed at the structural layer should not be fixed at the prompt layer. Too many tools: cut tools. Dirty context: add a gate. Runaway loop: add a budget.
- The prompt handles only what cannot be expressed structurally: domain preferences, tone, the soft edges of judgment calls.
- For every prompt rule you add, ask: "Six months from now, after a model swap, will this rule still be needed?" If the answer is mostly "no," it probably belonged in the harness all along.
The SWE-agent paper offers a positive reference point here: the key to their success-rate gains was not a better prompt but a redesign of the interface between the agent and the computer (the ACI) — making correct behavior easier and incorrect behavior harder (arXiv 2405.15793).
8. Demo-Driven Development
Symptoms
Capability is proven by "look, let me demo it" — three demos, two successes, screenshot the successful one. After launch, when users report bad cases, the team's fix process is: reproduce manually → tweak something → reproduce manually again → pray nothing new broke. No number anywhere can answer "is this week's version better or worse than last week's?"
Root cause
Agent behavior is probabilistic; a single demonstration carries almost no information. Without evals, every change is groping in the dark: prompt tweaks, model upgrades, adding or removing tools — each one might improve one cluster of tasks while collapsing another, and you have no instrument to see it happen.
The fix
- Build an eval set on day one, even if it's only 20 cases: real tasks + expected results + a checker that can score automatically (did the tests pass? does the diff match? another model as judge?).
- Hook evals into CI: every prompt change, model swap, or tool change triggers a run; watch three curves — success rate, average steps, average cost.
- The destination for bad cases is the eval set, not the chat log — turn every production failure into a regression case.
- This and Evaluation & Observability are two sides of one coin: traces show you a single run; evals show you the overall trend.
A pragmatic starting point
Don't wait for the "perfect eval framework." A JSONL file plus a script that prints a pass rate at the end already beats where most teams stand. Get numbers first; talk about the quality of the numbers after.
9. Blind to Cost & Latency
Symptoms
Acceptance testing is all green — until finance or a user comes knocking: $8 average per task, 40-second P50 first response; the support agent takes three planning steps, two tool calls, and 90 seconds to answer "what are your opening hours?"
Root cause
During development, only one metric was watched: success rate. But agent systems burn money multiplicatively with step count: every step re-sends (or re-summarizes) the whole context, so token consumption grows roughly as the square of the number of turns, and every step's latency lands directly on the user's wait time. One figure worth remembering from Anthropic's context engineering article: agents consume tokens at a scale far beyond ordinary chat, and multi-agent architectures amplify that further (Effective context engineering for AI agents).
The fix
- Put cost into the trace: every span records input/output tokens, model, unit price, and duration; roll up per task into the three metrics "cost / steps / duration," on the same dashboard as success rate.
- Tiered routing: route simple requests to a small model, or even deterministic code; save the big model for genuinely hard problems. Routing is one of the foundational patterns in Building effective agents and the best bang-for-buck optimization on this list.
- Cache everything you can: the system prompt and tool schemas are stable — use prompt caching; cache repeated retrieval results by key.
- Parallelize independent tool calls; don't serialize steps that could run in parallel.
- Set budget alerts (echoing #4): when a single task's cost crosses the threshold, prefer degrading to "partial result + explanation" over grinding all the way to the end.
10. Premature Multi-Agent
Symptoms
The single agent isn't even tuned yet, but the architecture diagram already shows five roles: Orchestrator, Researcher, Coder, Reviewer, Critic. At runtime, the five agents each see only their own context: the Coder doesn't know what the Researcher found, the Reviewer is criticizing code built on a wrong premise, and the Orchestrator dispatches the same task twice. The token bill is several times a single agent's, and the success rate is lower.
Root cause
Multi-agent splits the context, not the capability. Sub-agents don't share working memory, so all coordination happens through explicit message passing — and every handoff is a lossy compression. Cognition (the Devin team) makes the point sharply in Don't Build Multi-Agents (2025-06, Walden Yan): when parallel sub-agents each make their own decisions, the moment their implicit assumptions diverge, the merge stage produces conflicts that cannot be reconciled; their principle is "share context, share the full trace."
Example scenario
A team split their data-analysis agent into a "data-fetching agent" and a "charting agent." The data-fetching agent found one field was all nulls and switched to a different table — a decision that existed only in its trace. The charting agent then carefully laid out a chart against "the wrong table" and wrote an interpretation to go with it. To locate the divergence, a human had to replay both traces side by side; debugging cost doubled.
The fix
- First push single-agent + good context management to its ceiling: compaction, just-in-time retrieval, structured notes (see Memory Systems). A lot of "we need multiple agents" is really "our single agent's context is dirty."
- Genuine signals that multi-agent applies: the subtasks are parallelizable, loosely coupled, and independently verifiable (breadth-first research is the classic case). If those three don't hold, splitting only amplifies the coordination tax.
- The halfway option is the Subagents pattern: a sub-agent does focused work in a clean window and returns only a compressed conclusion to the main agent — you keep the benefit of context isolation while decision-making stays in one place.
- Read the counterargument before deciding: LangChain's How and when to build multi-agent systems (2025-06) responds to Cognition's post and draws the boundary of applicability fairly on both sides.
Mapping Anti-Patterns to Harness Components
When troubleshooting, use this table to work backwards to "which chapter should I read":
| Anti-pattern | Primary component |
|---|---|
| Framework-first | Overall architecture — see Anatomy of the Harness |
| Tool sprawl, swallowed errors, prompt-only tuning (partly) | Tool Design |
| Context hoarding, prompt-only tuning (partly) | Context Engineering |
| No termination conditions | The Agent Loop |
| Full-auto everything | Permissions, Safety & Human-in-the-Loop |
| Demo-driven development, blind to cost & latency | Evaluation & Observability |
| Premature multi-agent | Subagents |
Three Universal Defense Principles
Abstract all the anti-patterns up a level and there are only three root diseases — and only three defenses:
- Everything goes into the trace, and the trace goes into evals. Silence is the harness's number-one enemy. A run you can't see cannot be debugged, and an improvement you can't quantify is not an improvement.
- Write constraints into the structure, not into the prayers. Budgets, permissions, termination conditions, output gates — anything "the model shouldn't do" should be something the model cannot do, not something it was told not to do.
- Complexity is debt; decide how much to borrow by how fast the interest accrues. Frameworks, multi-agent, long prompts, big toolboxes — every one of them is borrowed complexity. Before borrowing, ask: is the interest it charges (debugging cost, context noise, coordination tax) below the value it delivers?
Further Reading
- Design Principles: the positive version of this page — anti-patterns tell you what not to do; design principles tell you what to do
- Build Your Own Harness: step through the shallow versions of these pitfalls yourself with a minimal implementation
- Model vs. Harness: why most of these pitfalls have little to do with model capability
- Classic Papers, Annotated: full reading guides for the papers cited here — ReAct, Lost in the Middle, SWE-agent, and more
References
- Anthropic — Building effective agents (2024-12-19): simple composable patterns beat complex frameworks; tool polish costs more than prompt tuning; routing and other foundational patterns
- Anthropic — Effective context engineering for AI agents (2025-09-29): context rot, attention budget, compaction, just-in-time retrieval
- Cognition — Don't Build Multi-Agents (2025-06-12): context fragmentation in multi-agent systems and the "share context" principle
- LangChain — How and when to build multi-agent systems (2025-06-16): a response to Cognition's post, on where multi-agent does and doesn't apply
- Liu et al. — Lost in the Middle: How Language Models Use Long Contexts (2023, arXiv 2307.03172): the U-shaped collapse of mid-context information utilization
- Yao et al. — ReAct: Synergizing Reasoning and Acting in Language Models (ICLR 2023, arXiv 2210.03629): the reasoning-acting interleaved agent loop paradigm
- Yang et al. — SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (NeurIPS 2024, arXiv 2405.15793): how interface/tool design significantly affects agent success rates
- How Many Tools Should an LLM Agent See? (arXiv 2605.24660): the negative effect of candidate tool count on selection accuracy
- RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation (2025, arXiv 2505.03275): retrieval-based tool selection for oversized registries
- QLCoder: A Query Synthesizer for Static Analysis (arXiv 2511.08462): an engineering retrospective recording that "a big toolbox causes selection chaos; a small one is more reliable"