Appearance
Agent Design Principles
The Agent field doesn't lack demos; it lacks systems that run stably in production. The good news: over the two-plus years from late 2024 to 2026, Anthropic, Princeton (the SWE-agent team), OpenAI, and a group of independent practitioners have published the pitfalls they hit, and these lessons converge—converge enough to be summarized in eight principles.
This page takes the eight principles apart one by one: where each comes from, what designs count as following it (positive examples), what counts as violating it (negative examples), and a set of check items you can lift straight into code review. They aren't eight parallel suggestions; they have an internal order: first decide whether to build it at all (1), then decide how to design the interfaces (2, 5), then make it run stably (3, 4, 6), and finally make it keep improving (7, 8).
Before diving in, read Anatomy of an Agent and the Agent Loop first—the discussion below assumes you know the basic runtime structure of an agent.
1. Simplicity first: workflow over agent, single agent over multi-agent
Source
The authoritative statement of this principle comes from Anthropic's December 2024 article Building Effective Agents:
When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.
The same article gives a distinction that remains the industry's standard answer: a workflow is a system where "LLMs and tools are orchestrated through predefined code paths," while an agent is a system where "the LLM dynamically directs its own process and tool usage." After observing a large number of customer projects, Anthropic's conclusion was: the most successful implementations were almost never built with complex frameworks, but assembled from simple, composable patterns; for many applications, a single well-optimized LLM call plus retrieval and few-shot examples is already enough.
Positive examples
- Invoice field extraction: one prompt + structured output—not even a workflow, but accuracy is on target and cost is minimal.
- Support-ticket triage: a routing workflow—classify first, then route into the corresponding specialized prompt. Fixed paths; predictable and testable.
- Code review: prompt chaining—first have the model list suspicious spots, pass a programmatic check (gate), then expand each one.
Negative examples
- A "translate + proofread" requirement split into 5 agents talking to each other, multiplying latency and token cost while errors propagate and amplify along the chain.
- A task with a completely fixed path (say "fetch → summarize → store") wrapped in an autonomous agent framework with reflection loops, where the framework's abstraction layers even hide the prompt during debugging—Anthropic explicitly warned about this: the extra abstraction of frameworks can obscure the underlying prompts and responses, making debugging harder.
Multi-agent is the exception, not the default
Anthropic's own multi-agent research system (the build retrospective published in June 2025) proves multi-agent does work for breadth-first, parallelizable research tasks—but that was the choice after validating that a single agent's context couldn't cope, not the starting point. The costs and applicability boundaries of multi-agent are covered in Multi-Agent Architectures.
Check items
- Can it be solved with a single LLM call (plus retrieval/examples)? If yes, stop there.
- Is the task path knowable in advance? Yes → workflow; no → only now consider an agent.
- For each added layer of complexity, can you articulate the measurable benefit it buys?
- If you're using a framework, can you read every prompt it sends out?
2. Design interfaces for the model, not for humans
Source
This principle comes from the SWE-agent paper (SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, arXiv: 2405.15793, NeurIPS 2024). The authors proposed the concept of the ACI (Agent-Computer Interface), mirroring HCI for humans. The paper's core finding: the same model, given an interface tailored for an LLM (LM-friendly commands, cleanly formatted feedback), improves its SWE-bench solve rate by an order-of-magnitude margin—interface design is itself a source of performance.
Anthropic's September 2025 Writing Effective Tools for Agents — with Agents brought the same idea down to the tool level: don't build thin wrappers over APIs; build "high-leverage tools"; tool names must be unambiguous; return meaningful, human-readable context instead of bare IDs; control output tokens. They even used agents to run evals and iterate on tool descriptions; after iteration, task completion time dropped by roughly 40% (a similar practice reported in the Anthropic multi-agent research system retrospective).
Positive examples
SWE-agent's edit command design is a textbook case. Rather than letting the model emit sed commands or diff patches directly (both formats models get wrong easily), it provides an edit command with line-number ranges, plus the convention that "only one file may be open at a time." Feedback is also curated: lint errors keep only the relevant lines, rather than dumping the whole terminal output at the model.
Negative examples
- Wrapping an internal REST API one-to-one as tools:
list_users+list_events+get_user_by_id… For "schedule a meeting with Zhang San tomorrow afternoon," the model must chain 4 calls, each one a chance to choose wrong. Anthropic's advice is to merge into a high-level tool likeschedule_event. - Tool descriptions written like human-facing API docs (the parameter names are the whole explanation), leaving the model to guess.
- Stuffing a 100k-line log file into a tool's return value.
The one-sentence test (Anthropic's words): if an engineer can't tell which tool to use from its name and description alone, neither can the model. For the full treatment of tool design, see Tools & MCP.
Check items
- Show each tool's name and description to a new colleague unfamiliar with the system—can they state its use case accurately?
- Does the tool return "the information the model needs next," or "whatever happened to be in the database"?
- Do error returns include actionable hints (e.g. "parameter date must be ISO format, received 'tomorrow'"), not bare stack traces?
- Have you iterated on tool descriptions using real traces, or written them once and never touched them again?
3. Make the agent observable and explainable: announced actions, transparent trajectories
Source
Agents are non-deterministic systems: the same input can take entirely different paths across two runs. Without a trace, a failed agent leaves you with nothing but "no idea why it did that." Products like Claude Code make this the default at the interaction level—each action is "announced" before execution (what command I'll run, which file I'll change), results are shown after, and the whole trajectory is replayable. At the infrastructure layer, OpenTelemetry's GenAI semantic conventions (the gen_ai.* namespace) became the vendor-neutral standard for LLM/Agent telemetry across 2025-2026: span types like invoke_agent and execute_tool, and attributes like gen_ai.request.model, keep your instrumentation decoupled from your observability backend.
Positive examples
- Every LLM call and every tool execution is a span, with input, output, token usage, and latency; the whole trace is replayable.
- Destructive actions are clearly announced in the UI and wait for confirmation before executing (this pairs with principle 6's tiered permissions).
- Production emits spans per the OTel GenAI conventions, so switching observability platforms (Langfuse, LangSmith, self-hosted Grafana) requires no re-instrumentation.
Negative examples
- Logging only the final answer; the 20 tool calls in between are a black box.
- Interpreting "observability" as "an accuracy dashboard"—the lights are green while every answer is wrong. Traces are for reading line by line, not just for aggregate metrics.
Read 10 traces yourself every week
Hamel Husain's "look at your data" applies just as much to agents: aggregate metrics tell you that something happened; only reading trajectories one by one tells you why. Put "read traces weekly" on your calendar—it matters more than buying any observability platform. The full observability stack is covered in Observability.
Check items
- For any failed run, can you pull the complete trajectory and locate the failing step within 5 minutes?
- Does the trajectory record both "what the model saw" and "what the tool actually did"?
- For end users, do long agent tasks announce actions and show progress?
- Is token usage and cost attributable per trace (which step burned the money)?
4. Failure is the norm: design for recovery
Source
Traditional software assumes dependencies are mostly reliable; the agent world is the opposite: models misread, tools time out, pages change, APIs rate-limit. A counterintuitive finding in the SWE-agent paper: telling the model "your last command failed" is more valuable than trying to prevent failure—models are quite good at self-correcting from clear error feedback. Anthropic's tool design guide makes the same point: error responses should "steer" the agent toward less token-hungry, more correct behavior (e.g. suggesting pagination or filtering).
Positive examples
python
# Tool execution layer with recovery strategy: three tiers—retry, degrade, escalate
def execute_with_recovery(tool_call, max_retries=2):
for attempt in range(max_retries + 1):
try:
return run_tool(tool_call)
except TransientError as e: # timeout, rate limit: worth retrying
if attempt < max_retries:
backoff_sleep(attempt)
continue
return tool_error(f"failed after {max_retries} retries: {e}. "
f"Suggest switching tools or narrowing the request.")
except PermanentError as e: # bad arguments, missing permission: retrying is pointless
return tool_error(f"call rejected: {e}. Check the arguments and retry, "
f"or try another approach.")The key design: errors aren't thrown upward to blow up the flow; they're formatted and handed back to the model, letting it decide the next move—change the arguments, switch tools, or admit this road is blocked. Only true dead ends (beyond the model too, say the target system is down) escalate to humans, see Human-in-the-loop.
Negative examples
- Using the same retry policy for 4xx (bad arguments) and 5xx (service failure)—waiting through three retries even when the arguments are wrong.
- Returning an empty string when a tool fails, so the model reads it as "no results found" and keeps reasoning on that basis—silently swallowed errors are the worst kind of failure.
- No max-step/budget circuit breaker, and the agent burns hundreds of dollars of tokens down one dead-end direction.
Check items
- Have you separated retryable from non-retryable errors? Do retries back off and cap out?
- For each kind of tool failure, is the information returned enough for the model to act differently?
- Is there a global circuit breaker (max steps, max tokens, max wall-clock time)?
- For irreversible operations, can failures roll back or at least leave an audit record?
- Does a "total failure" path exist—can the agent say "I can't do this" instead of hallucinating an answer?
5. Context is a scarce resource: every token must earn its place
Source
Anthropic's September 2025 Effective Context Engineering for AI Agents opens with the thesis: context is the key but finite resource. It's not just about the context window not fitting—even when it fits, attention and recall quality degrade as context grows (the so-called context rot), and token cost climbs linearly. The 2026 reality: mainstream flagship models all support 1M-token-class windows, but "can stuff it in" and "still works stuffed in" are different things, and pricing structures (some models charge a premium for very long inputs) remind you that context isn't free.
Positive examples
- Compaction: as a long task nears the window limit, summarize the early trajectory into structured notes, keeping key decisions and open questions.
- Structured notes: the agent writes "confirmed facts" and "approaches that failed" to external files and reads them back via tools when needed, rather than keeping them in the prompt forever.
- Sub-agent isolation: exploratory tasks (say, "find the relevant code among these 50 files") go to a sub-agent; the main agent only receives the distilled conclusion—Anthropic's multi-agent research system solves context bloat exactly this way.
- Tool output trimming: truncate, paginate, and filter by default; hand the "how much to fetch" decision to the model.
Negative examples
- Concatenating the entire conversation history plus all retrieval results indiscriminately into every turn's prompt.
- Stuffing 30 edge-case rules into the system prompt "just to be safe," and now the model starts fumbling even on simple cases—every line in context competes for attention.
This methodology deserves a deep read on its own, see Context Engineering.
Check items
- Do you know the token composition of a typical task (how much for system prompt / history / tool output)?
- What's your compaction strategy near the window limit, and have you validated the success rate after compaction?
- Do tool return values have truncation and pagination?
- For each instruction in the system prompt: would evals drop if you deleted it? If not, delete it.
6. Autonomy matches risk: tiered permissions
Source
"How much autonomy to give an agent" shouldn't be a global switch; it should be tiered by each action's reversibility and blast radius. Claude Code's permission modes are the industry's reference implementation: by default, reads are free and writes need confirmation; acceptEdits loosens editing; bypassPermissions (community nickname: YOLO mode) opens almost everything, with official guidance to use it only in isolated environments (containers, VMs without sensitive data). The underlying idea is simple: the degree of autonomy should be proportional to the cost of a mistake, and to how isolated the environment is.
Positive examples
A three-tier permission design:
python
from enum import Enum
class RiskLevel(Enum):
READ_ONLY = 0 # read files, search, query: pass through directly
REVERSIBLE = 1 # write files, change config: can run automatically, but log to the audit trail
IRREVERSIBLE = 2 # send email, delete data, trigger payment: must have human confirmation
def authorize(action) -> bool:
if action.risk == RiskLevel.READ_ONLY:
return True
if action.risk == RiskLevel.REVERSIBLE:
audit_log(action) # leave a record, rollback possible after the fact
return True
return ask_human(action) # escalate to a human, with full context attachedTiered permissions are just one layer of defense in depth; prompt injection and other agent-specific attack surfaces are covered in Security.
Negative examples
- Running
--dangerously-skip-permissionson your primary dev machine from day one, and the agent executesrm -rfafter being injected by malicious web content. - Conversely, requiring human confirmation for every action—people get numb and start blindly clicking "allow," the confirmation mechanism becomes theater, and it's ten times slower than full automation. Permission design fails in two directions; overtightening is as common as over-loosening.
Check items
- Have all tools been tiered by risk? Are irreversible operations forced through human confirmation?
- Are high-permission modes only enabled in isolated environments?
- Is the audit log independent of the agent (the agent can't alter its own operation records)?
- Does the human-confirmation UI show enough context (full command, target files, blast radius), rather than just "allow this operation?"
7. Evals first: optimization without evals is blind
Source
Hamel Husain's Your AI Product Needs Evals (March 2024, still the most-cited practical guide to eval methodology) makes the core claim: what separates successful LLM product teams from unsuccessful ones isn't fancy techniques, but whether they've built the iterative loop from looking at data, to error analysis, to writing evals. Read 100+ real traces first, categorize the failure modes, write evals for the most frequent failures, and only then change the prompt or the architecture—the order can't be reversed.
Positive examples
- Before changing a prompt, there's a 50-200 item eval set (from real failure cases); after the change, run it and let the numbers speak.
- Error-analysis-driven: discovering that 40% of failures are "wrong tool argument format," so fixing the tool description first (principle 2), rather than switching to a pricier model.
- Layered evals: deterministic unit tests (format, schema) + LLM-as-judge (quality) + end-to-end task success rate, each covering its own ground.
Negative examples
- "I feel like the new prompt is better"—iterating by vibes, and two weeks later the system has quietly degraded and nobody noticed.
- Only benchmark scores (SWE-bench, AgentBench, and the like) with no eval for your own business. Public benchmarks measure model selection; they can't measure whether your system is reliable on your traffic.
For the full eval-system build-out, see Agent Evaluation and Evals in Practice.
Check items
- Is there an eval set that runs on every significant change? Does it come from real failure cases?
- Are failure modes categorized and counted (which kind of error dominates)?
- Before switching models, changing prompts, or changing tools, can you estimate the impact and verify it afterwards?
- Has the eval itself been validated (how many LLM-judge verdicts have been sampled for human review)?
8. Simplify as models evolve: yesterday's necessary scaffolding may be today's dead weight
Source
This principle's roots are Rich Sutton's 2019 The Bitter Lesson: the biggest lesson of 70 years of AI research is that general methods leveraging computation ultimately beat hand-crafted methods that encode human knowledge. Boris Cherny, creator of Claude Code, has said in interviews that the team's workspace wall has a framed copy of The Bitter Lesson, and their translation is: never bet against the model. You can write scaffolding (all the code beyond the model) that improves a scenario by 10-20%, but the next model generation may do it on its own, and your scaffolding becomes the constraint—the LangChain team publicly recounted a similar experience in 2025: rigid agent structures designed for weak models became bottlenecks in the strong-model era.
Engineering implication: scaffolding is rented capability, not an asset. Every scaffolding component should carry the expectation of "remove it once the model gets stronger."
Positive examples
- A "forced step-by-step planning" template written for weak models gets retired in the reasoning-model era, letting the model plan itself—evals confirm it doesn't drop, and may even improve.
- A complex edit protocol designed to prevent models from writing malformed diffs gets progressively simplified as model capability rises (the SWE-agent team's own mini-SWE-agent is an experiment in stripping the ACI to the bare minimum).
- Every scaffolding component sits behind a feature flag, so on a model swap you can turn them off one by one and validate with evals.
Negative examples
- Using, verbatim, a 2023-era "chain-of-thought incantation + 12 behavioral rules" prompt on a 2026 model; the system prompt is longer than the task itself, and the model is shackled by outdated instructions.
- Scaffolding code with no switches and no tests—when you want to remove it, you can't tell what would break.
Every time you switch models, first ask "what can I delete"
The first item on a model-upgrade checklist shouldn't be "what new capabilities can I add," but "which scaffolding can retire." A model upgrade where you get to delete code is a real upgrade—after deleting, run principle 7's evals and confirm with numbers.
Check items
- Can every scaffolding component be described in one sentence as "which model weakness it patches"?
- Is the scaffolding behind a feature flag or config, so you can switch it off quickly and run a comparison eval?
- The last time you swapped models, did you delete anything? If you've never deleted anything, you're probably carrying a stack of expired baggage.
9. The Eight Principles at a Glance
| # | Principle | In one sentence | Source |
|---|---|---|---|
| 1 | Simplicity first | Single call over workflow, workflow over agent, single agent over multi-agent | Anthropic, Building Effective Agents |
| 2 | Design interfaces for the model | Tools are high-leverage, unambiguous, token-efficient ACIs, not thin API wrappers | SWE-agent (arXiv: 2405.15793); Anthropic tool guide |
| 3 | Observable and explainable | Announce actions, replayable traces, standardized spans, read traces yourself weekly | OTel GenAI conventions; product practice like Claude Code |
| 4 | Design for recovery | Format errors back to the model, tiered retries, global circuit breakers, allow saying "can't do it" | SWE-agent; Anthropic tool guide |
| 5 | Context is scarce | Compaction, notes, sub-agent isolation, tool output trimming | Anthropic, Effective Context Engineering |
| 6 | Autonomy matches risk | Tier grants by reversibility; high permissions only in isolated environments | Claude Code permission modes |
| 7 | Evals first | Look at data and do error analysis first, then write evals, then change the system | Hamel Husain |
| 8 | Simplify as models evolve | Scaffolding is rented; on a model swap, first ask what can be deleted | Sutton, The Bitter Lesson; Boris Cherny |
Put together, the eight principles are really one sentence: spend complexity where it pays, leave simplicity to the model, and keep control for yourself. To apply these principles in a complete build, start with Build Your First Agent by Hand, then self-check against Common Pitfalls when you're done.
References
- Building Effective Agents — Anthropic — source of principle 1 and the page's overall tone; the standard workflow/agent distinction.
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (arXiv: 2405.15793) — the original paper for the ACI concept, NeurIPS 2024.
- Writing Effective Tools for Agents — with Agents — Anthropic's tool design guide (2025-09); high-leverage tools and agent-driven tool iteration.
- Effective Context Engineering for AI Agents — the systematic treatment of context as a finite resource (2025-09).
- How We Built Our Multi-Agent Research System — an engineering retrospective on when multi-agent is worth it (2025-06).
- Your AI Product Needs Evals — Hamel Husain — the practical guide to evals-first and error-analysis methodology.
- The Bitter Lesson — Rich Sutton — the theoretical root of principle 8 (2019-03).
- OpenTelemetry: Inside the LLM Call — GenAI Observability — the official introduction to the
gen_ai.*semantic conventions, the vendor-neutral standard for agent telemetry.