Skip to content

Tools & MCP

At a glance Tools are the interface between a model and its environment, and the most underrated part of harness design: this article breaks down what deserves to be a tool, how schemas and return values shape model behavior, the choice-overload problem of having too many tools, what the MCP protocol does and doesn't solve, the tradeoffs of general-purpose computer use tools, and the peculiar form of prompt-only tools like TodoWrite.

Tools & MCP ​

Tools answer one question: how does the model reach the world beyond its context?

A model on its own does exactly one thing: tokens in, tokens out. Everything else it appears to do — reading files, running tests, querying databases, sending messages — happens because the harness translates a structured action request into a real function call, then translates the result back into text and feeds it into the context. In other words, tools are the interface between the model and its environment, and interface design comes with a rule of thumb that countless products have validated: the quality of the interface sets the ceiling on what its user can achieve.

The same rule holds for agents, except the "user" is now the model. The SWE-agent paper (arXiv:2405.15793, NeurIPS 2024) crystallizes this into a concept — the Agent-Computer Interface (ACI): just as a good UI determines how well a person can operate software, the commands and feedback formats you give a model determine how well it can operate a computer. The SWE-agent team found that simply changing the edit command from "a free-text patch" to "a line-range replace," paired with compact error feedback, was enough to significantly improve SWE-bench scores — the model didn't change, the prompt didn't change, only the tool interface did.

Hence this article's central claim: tool design is the most underrated part of harness design. People love debating which model to use and which agent framework to adopt, yet rarely stop to answer a more basic question: how are the names, parameters, and return text of these tools shaping the model's behavior at every single step? Anthropic's engineering blog post Writing effective tools for agents offers a precise definition that makes a good starting point:

A tool is a new kind of software — not a contract between deterministic systems, but a contract between a deterministic system and a non-deterministic agent. When you write an API for getWeather("NYC"), the caller is another program; an agent might call it, might answer from memory, or might turn around and ask the user, "Which city are you in?"

Understanding what makes this contract special is the prerequisite for understanding every tool design decision that follows.

What deserves to be a tool: not every API needs wrapping ​

The most common mistake is wrapping existing API endpoints one-to-one as tools: one list_contacts endpoint becomes one list_contacts tool; three endpoints like get_user, list_transactions, and list_notes become three tools. That's the natural order of things in traditional software integration, but in agent scenarios it's an anti-pattern — because the agent's "memory" is an expensive, finite context window, not cheap heap space.

Anthropic's blog post offers a few before-and-after examples:

  • Not list_contacts (returns every contact, forcing the model to page through them), but search_contacts (returns only the relevant few);
  • Not list_users + list_events + create_event (the model chains the three steps itself), but schedule_event (one tool checks availability and creates the event internally);
  • Not get_customer_by_id + list_transactions + list_notes, but get_customer_context (returns the customer's full relevant context in one shot).

The pattern is clear: design tools around workflows, not around API endpoints. Every merge removes one round trip of intermediate results through the context, and one more chance for the model to garble data while "transcribing" it between steps.

But this doesn't mean coarser tools are always better. The mistake in the other direction is the universal Swiss Army knife: a single execute_action tool that takes a free-text instruction parameter. That pushes all the interface design work back onto the model — effectively no tool at all. A rule of thumb for what deserves to be a tool:

  • High-frequency, multi-step, deterministic workflows get merged into a single tool (like schedule_event);
  • Branch points that require the model's judgment in the moment stay as separate tools — don't make the decision for it up front;
  • Things the model can already compose with shell + code don't need a dedicated tool — Claude Code ships only "general primitives" like Read/Write/Grep/Bash and leaves the composition to the model, a deliberate act of restraint (see the Claude Code case study);
  • Purely informational capabilities with no side effects on the environment belong in Skills (instructions) rather than tools (runtime), to keep the tool list from bloating.

Schema design: interface docs written for the model ​

A tool's schema — its name, description, and parameter definitions — is the model's only basis for choosing and using the tool. It never touches a database or a compiler; it goes straight into the prompt. Schema design is therefore context engineering in disguise: every word you write shapes the model's behavior distribution.

Naming is routing ​

The tool name is the first signal the model uses for routing decisions. When an agent has dozens of MCP servers and hundreds of tools attached at once, overlapping or vague names directly cause miscalls. Anthropic's field-tested advice is namespacing: group by service prefix (asana_search vs jira_search), layer by resource (asana_projects_search vs asana_users_search), so the name itself carries the information "when should you use me." They even found that prefix- versus suffix-style naming has a "non-trivial impact" on tool-calling accuracy, and that different models reach different conclusions — which means your naming scheme belongs in your evals, not in gut feel.

Descriptions are prompt engineering, not documentation ​

A tool description isn't reference documentation for humans — it's a behavioral instruction aimed at the model. Anthropic's blog has a textbook example: after their web search tool launched, they noticed Claude kept appending a gratuitous 2025 to the query parameter (presumably a bias inherited from training data, where recency queries come with a year), polluting the search results. The fix wasn't a code change and it wasn't fine-tuning — it was editing one line of the tool description to steer the model back on track.

A few actionable rules fall out of this:

  • State clearly when to use the tool and when not to, not just what it is;
  • Give positive and negative examples for parameters that are easy to fill in wrong (what query should and shouldn't look like);
  • Complex, domain-specific usage discipline works better in the system prompt than crammed into a tool description (Anthropic verified this in the think tool's τ-bench experiments, covered later in this article);
  • Constrain parameters at the schema level (enums, required fields, format notes) to eliminate half of the possible mistakes structurally — the model doesn't need persuading; it needs zero ambiguity.

Identifiers should be human (and model) friendly ​

One easily overlooked detail: models handle natural-language identifiers far more reliably than meaningless strings. Anthropic found that resolving arbitrary UUIDs in tool results into meaningful names (even a zero-based ordinal) measurably reduces hallucinations in retrieval tasks — it's easy for a model to miscopy a character when quoting uuid: 7f3a..., but quoting alice-smith almost never goes wrong. Similarly, return fields should lead with information downstream decisions can use directly, like name and file_type, over low-level fields only a program cares about, like mime_type or 256px_image_url.

Return-value design: what a tool says is also a prompt ​

Tool return values and errors enter the context as messages with the tool role, becoming the input for the model's next round of reasoning. So the audience for return-value design isn't the developer — it's the model.

Error messages are hints for the model ​

Traditional API errors are written for the calling programmer: {"error": 400, "code": "INVALID_PARAM"}. To a model, an error like that carries almost zero information — it'll just retry with a different parameter value, or give up entirely. A good tool error should read like a colleague walking the model through the fix:

text
Bad error:    Error: invalid parameter
Good error:   Error: `path` must be a repo-relative path (e.g. "src/main.ts"),
              got the absolute path "/home/user/proj/src/main.ts".
              Tip: use the ListFiles tool to inspect the directory structure first.

The good version does three things at once: pinpoints what went wrong, shows the correct format, and recommends the next move. Most of the moments where a model gets stuck trace back to a tool that told it "you're wrong" without telling it "what right looks like."

Token budgets and truncation ​

Tool returns are one of the leading causes of context bloat. A grep that returns five thousand matching lines, or an API call that returns an entire JSON document, can push critical instructions out of the model's effective attention (the "context rot" that context engineering fights against). Standard engineering practice:

  • Cap by default: Claude Code truncates tool output at 25,000 tokens by default (the figure disclosed in Anthropic's blog post);
  • Truncation must come with guidance: never cut silently. A good truncation message looks like "Results truncated at 100 items. Use filters or pagination to narrow results." — telling the model both that results were cut and what to do next, nudging it toward many small, precise queries instead of one giant catch-all;
  • Provide pagination/filtering/range parameters: let the model control how much comes back, and expose switches like response_format: concise | detailed where appropriate.

The full anatomy of a tool call ​

Putting the principles together, here's the complete lifecycle of a tool call inside a harness:

text
┌────────────────── Anatomy of a tool call ───────────────────┐
│                                                             │
│  ① schema enters the system prompt                          │
│     (name/description/params — the model's basis for choosing) │
│              │                                              │
│              ▼                                              │
│  ② the model emits a structured call {name, args}           │
│              │                                              │
│              ▼                                              │
│  ③ harness validation + permission gate — see /components/permissions │
│              │                                              │
│              ▼                                              │
│  ④ real execution (files/network/database)                  │
│              │                                              │
│              ▼                                              │
│  ⑤ results feed back into context — high-signal fields, error guidance, truncation │
│              │                                              │
│              ▼                                              │
│  ⑥ the model continues to the next turn of the agent loop   │
└─────────────────────────────────────────────────────────────┘

Note that steps ① and ⑤ are both "text going into context" — both ends of a tool are essentially prompt design, and only ④ is classic software engineering. That's why tool design demands backend thinking and prompt thinking at the same time.

Tool count and choice overload ​

More tools is not better. There are two independent costs here.

The first is static overhead: schemas eat tokens. Every tool you attach keeps its definition resident in the system prompt. With dozens of MCP servers and hundreds of tools, the model has to chew through hundreds of thousands of tokens of tool definitions before it even sees the user's request (the order of magnitude Anthropic describes in Code execution with MCP) — responses get slower, costs climb, and the task's own context gets squeezed out.

The second is behavioral overhead: selection accuracy drops as candidates multiply. The model is making a classification decision over a pile of candidates; the more candidates there are and the more alike they look, the higher the odds of picking the wrong tool, filling in wrong parameters, or failing to call a tool when it should. Anthropic puts it bluntly: too many tools, or tools with overlapping functionality, "distract the agent and pull it away from efficient strategies." Cloudflare's Code Mode post gives an even more pointed explanation: the special token format for tool calling comes from only a small amount of synthetic training data, while code comes from millions of real open-source projects — having the model write code that calls APIs plays much closer to its strengths within the training distribution than having it call tools directly.

The remedies have converged into a standard playbook:

SolutionMechanismNotable example
Merge toolsConverge around workflows to shrink the candidate poolget_customer_context replaces three lookup tools
NamespacingGroup by prefix to cut down misselectionasana_search / jira_search
On-demand loadingHand over a search_tools retrieval tool first; load a tool's full definition only when it's usedAnthropic's "progressive disclosure"
Code executionTurn tools into a code API; the model writes code that calls them and filters data in the execution environmentAnthropic / Cloudflare Code Mode

The numbers on that last route are startling: after Anthropic rewrote multi-tool workflows like "pull meeting notes from Google Drive and write them into Salesforce" as code execution, token consumption dropped from 150,000 to 2,000 (-98.7%) — intermediate results no longer pass through the model's context twice, and loops, conditionals, and retries all happen inside the execution environment. The price is a secure sandboxed execution environment, a new piece of infrastructure to carry; the tradeoffs are covered in Permissions & human-in-the-loop.

A meta-rule

The right unit for tool count isn't "a number of tools" — it's the decision cost the model pays to pick the right one. Ten tools with crisp responsibilities and orthogonal names beat fifty with blurry boundaries; and five hundred tools usually shouldn't all go into context at all — the right move is to add a retrieval layer or code execution.

MCP: the standardization layer for tooling ​

Once you've absorbed the finer points of tool design, MCP (Model Context Protocol) becomes much easier to place: what it actually solves, and what it leaves open.

What it is ​

MCP is the protocol Anthropic open-sourced on November 25, 2024, aiming to set a single standard for "AI applications connecting to external tools and data." The architecture is the classic host/client/server split: clients inside the harness (the host) open sessions with MCP servers, and servers expose three kinds of primitives over JSON-RPC 2.0 — tools (callable functions), resources (readable data), and prompts (reusable instruction templates). Transport can be local stdio or remote HTTP.

text
┌────────────── Host (your harness) ──────────────┐
│                                                 │
│   Agent Loop                                    │
│      │                                          │
│      ▼                                          │
│   ┌─────────┐  JSON-RPC   ┌───────────────┐
│   │ Client  │ ══════════> │ MCP Server A  │ ──> local filesystem
│   │         │  (stdio)    │ (tools/resources) │
│   │         │ ══════════> │ MCP Server B  │ ──> GitHub API
│   │         │  (HTTP)     │               │
│   │         │ ══════════> │ MCP Server C  │ ──> database
│                           └───────────────┘     │
└─────────────────────────────────────────────────┘
        One protocol, N servers, plug and play

What it solves is the classic M×N integration problem: M agent applications × N external systems means M×N bespoke connectors without a protocol; with one, each side implements once and complexity drops to M+N. The community often reaches for the USB-C or LSP analogy (Language Server Protocol unified the interface between editors and language backends; MCP unifies the interface between agents and tools) — and that analogy is broadly accurate.

Adoption has been remarkably fast: in March 2025, OpenAI announced support for MCP across its product line (the two biggest rival labs backing the same protocol was the moment that made it the de facto standard); on December 9, 2025, Anthropic donated MCP to the Agentic AI Foundation (AAIF) under the Linux Foundation, where it sits alongside Block's goose and OpenAI's AGENTS.md as one of the three founding projects, with an official count of more than 10,000 published MCP servers. Governance moving from a single vendor to a neutral foundation essentially locked in its status as shared infrastructure.

What it doesn't solve ​

MCP standardizes the connection, not the quality. This is the distinction most easily blurred when judging MCP:

  • It says nothing about whether tools are well designed. An MCP server that wraps 50 API endpoints as-is is still bad tool design — none of the principles in the preceding sections get enforced for you. The protocol just makes bad tools easier to plug in.
  • It amplifies context bloat. In the default mode, the client pours every tool definition from every server into the context in one go. The more you attach, the worse the static and behavioral overheads from the previous section get — which is precisely why patterns like code execution and on-demand loading exist.
  • Its trust model is thin. Tool descriptions are instructions to the model, and an MCP server is third-party code that can carry malicious instructions. Invariant Labs disclosed tool poisoning attacks in April 2025: injection instructions hidden in tool descriptions (in the parts invisible to the user interface) that manipulate the agent into exfiltrating data; the same team later demonstrated rug pulls — servers quietly modifying tool definitions after the user has approved them. Installing an MCP server of unknown provenance is roughly equivalent to welcoming an untrusted prompt and a set of executable permissions into your harness at the same time. For the matching permission-gate design, see Permissions & human-in-the-loop.
  • Permissions, auditing, billing, and lifecycle management are all still rudimentary at the protocol layer; production deployments have to add this layer themselves (gateways, proxies, approval flows).

An honest assessment

MCP's value is real — it eliminates duplicate integration work and lets the tool ecosystem evolve independently. But "hooking up an MCP server" isn't the finish line of capability building; it's the starting line: once things are connected, whether you picked the right tools, whether the results are well designed, and whether permissions are tight enough remain entirely the harness's job. A richer ecosystem turns tool governance (which to adopt, how to trim, how to prevent injection) into the new core problem.

General interfaces vs. specialized tools: the computer use tradeoff ​

At the other end of the tool spectrum sit "general-purpose tools": no specialized interface at all — hand the model a screenshot plus a set of mouse and keyboard actions (click coordinates, type text) and let it operate the GUI the way a human does. Anthropic's computer use (released October 22, 2024, as a public-beta capability of Claude 3.5 Sonnet) and OpenAI's Operator (released January 2025) are the flagships of this route; open-source browser use projects are its variant in the browser domain.

The appeal is obvious: zero integration cost. No waiting for any software to expose an API, no writing MCP servers — as long as it has an interface, the agent can use it. For the long tail of legacy and internal software, it's the only realistic path to integration.

But the costs are just as structural:

DimensionGeneral interface (screenshot + click)Specialized tools (API / MCP)
Integration costPractically zeroRequires development or an off-the-shelf server
Per-step overheadA screenshot alone can burn a thousand tokens, one step at a timeOne call completes multiple steps
ReliabilityFragile: minor layout shifts, pop-ups, and loading delays scramble coordinatesDeterministic; verifiable and retryable
Auditability / interceptabilityThe action is "click (382, 511)" — ambiguous semantics, hard to reviewThe action is transfer(amount=100) — clear semantics, easy to gate
CoverageAny software with a GUIOnly systems that expose an interface

The industry's converged answer isn't either/or — it's layering: specialized tools as the main road, general interfaces as the safety net. Wherever an API exists, never have the model click pixels — not because it can't, but because every step pays for "generality" in tokens, latency, and failure rate. Two easily confused senses of "general" deserve untangling: Claude Code's Read/Bash are general primitives (atomic operations with clear semantics and deterministic execution), while computer use is a general interface (guessing semantics on someone else's UI). The former is tool design at its best; the latter is the floor that guarantees integration is always possible. The two can't substitute for each other.

Prompt-only tools: the zero-runtime-logic form ​

Finally, there's a form that's easy to overlook yet best embodies the "tool as interface" idea: the implementation does nothing at all — all of the effect comes from the mere act of the model calling it.

Two canonical examples:

TodoWrite. Claude Code's task-list tool has zero runtime logic — the harness doesn't check it, doesn't enforce it; it simply keeps the list the model wrote in context, where it serves as a visible anchor for every round of decisions. What makes it work is cognitive offloading: the act of writing the list forces the model to structure the task, and the list's continued presence counters goal drift on long-horizon tasks. This one tool reduces a planning problem to a tool-calling problem; the full breakdown is in Planning & task decomposition.

The think tool. Anthropic's think tool, released in March 2025, is even purer: its description says outright, "use this tool to think. It doesn't fetch new information, it doesn't modify the database — it just logs your thoughts." On evals like τ-bench, which demand adherence to complex policies and multi-step decisions, adding this empty tool measurably improves performance; on SWE-bench, giving Claude 3.5 Sonnet a think tool with a custom description produced the then-best score of 0.623. The mechanism is the same as TodoWrite's: give "stop and think" an explicit action slot, and the model will actually stop and think at the critical junctures.

The lesson for harness design runs deep: a toolset isn't just an inventory of "what the model can do" — it's scaffolding for model behavior. For whatever cognitive action you want the model to take at a given point (planning, reflection, progress reporting), you can design a corresponding prompt-only tool, then use system-prompt discipline to turn it into a fixed rhythm of the loop. Progressive disclosure in Skills and the structured-output conventions in the agent loop are the same idea applied at different layers.

Practical checklist ​

If you're designing the tool layer for your own agent, check these in order:

  1. List workflows first, then define tools. Derive the tool set backward from eval tasks (how real users will actually use this agent); working forward from API documentation is off-limits.
  2. Ask of every tool: can the model already compose this from existing tools? If yes, don't build it; if composing it is too roundabout, build a merged tool.
  3. Write schemas as prompts. Orthogonal names, descriptions that cover when to use the tool and counter-examples, parameters narrowed with enums and required fields.
  4. Write error messages for the model. Every error answers three questions: what went wrong, what the correct format is, and what to try next.
  5. Budget the return values. Default truncation + truncation guidance + pagination/filter parameters; high-signal fields first, and resolve UUIDs into names wherever possible.
  6. Past 30 tools, you need a loading strategy. Namespacing → on-demand loading via search_tools → code execution, escalating level by level.
  7. Review MCP servers as third-party code. Read tool descriptions in full (including the parts the UI doesn't show), pin versions, least privilege; see Common pitfalls.
  8. Write evals for your tools. Anthropic's methodology: prototype → eval tasks drawn from real scenarios → read failure traces and revise descriptions/schemas → a held-out test set to guard against overfitting. Tools can be optimized in an eval-driven loop; don't write them once and walk away.
  9. Consider whether you need a prompt-only tool. If the agent drifts on long tasks or skips thinking at critical points, TodoWrite / think is a nearly zero-cost intervention.

Further reading ​

  • Agent loop — the main loop structure that tool calling is embedded in
  • Context engineering — tool schemas and return values are both context content to manage
  • Planning & task decomposition — the full mechanism breakdown of TodoWrite, the prompt-only tool
  • Permissions & human-in-the-loop — where permission gates land on tools, and the countermeasure layer for MCP security problems
  • Skills — the division of labor that puts knowledge-style capabilities in instructions rather than tools
  • Observability — tool-calling traces as the first place to look when debugging agent behavior
  • Claude Code case study — the reference example of a "general primitives + prompt-only tools" tool philosophy
  • SWE-agent case study — where the ACI concept originated, and direct evidence that tool interfaces determine scores
  • Core papers — a guided tour of the original literature: ReAct, SWE-agent, and more

References ​