Skip to content

Context Engineering

At a glance Context is an agent's scarcest resource, and the harness's core job is deciding what the model sees at every step. Anatomy of a real context window, the four failure modes (poisoning, distraction, confusion, clash), compression and truncation strategies, KV cache hit-rate economics, and the trade-off between RAG and just-in-time retrieval.

Context Engineering ​

Of all the harness's responsibilities, context engineering is the most central one: the quality of every decision an agent makes isn't capped by the model — it's capped by the tokens the model sees at that moment. In Effective context engineering for AI agents, Anthropic offers a precise definition: context engineering is the entire discipline of curating and maintaining the set of tokens that enters the context window during LLM inference — with the goal of finding, for every step, the smallest, highest-signal token set that maximizes the probability of the desired behavior.

Why call it "engineering"? Because it isn't a one-shot task that ends once you've written a good prompt. Every turn of the agent loop can produce new tool results, new observations, new errors — candidate information keeps expanding, while the window is fixed and attention is finite. What gets in, what gets out, what gets compressed, what gets offloaded: at every step, these are questions the harness has to answer.

Context is the scarcest resource ​

"Scarce" has a hard meaning and a soft one.

The hard constraint: the window has a physical ceiling. Even with a 1M-token window, the tool trajectory of a long task (dozens of rounds of shell output, file contents, search results) will fill it up. Once it's full you either truncate or compress — both are lossy operations.

The soft constraint: the attention budget runs out before the window does. Citing needle-in-a-haystack-style benchmarks, Anthropic describes "context rot": as token count grows, the model's ability to accurately recall information in its context steadily degrades — not a cliff, but a slope. The architectural reason is straightforward: in a transformer, n tokens produce n² pairwise attention relationships, so the longer the context, the thinner each relationship's share of the attention budget. And because short sequences vastly outnumber long ones in training data, models simply have less "experience" with very long-range dependencies.

So the right mental model isn't "how much fits in the window" but: every token you add to the context spends a finite attention budget. Every context-engineering technique is, at bottom, a trade-off made under that budget constraint.

What's actually in the context ​

Before discussing how to manage it, dissect a real agent request and see what the context actually holds:

text
┌─────────────────────────── Context of one LLM call ─────────────────────────┐
│                                                                             │
│  ① System prompt                                                            │
│     Behavior rules, output format, tool discipline, safety boundaries       │
│                                                  ← stable, fixed overhead   │
│                                                                             │
│  ② Tool schemas                                                             │
│     Name/description/parameter definitions for every tool,                  │
│     serialized at the head of the context         ← more tools, more bloat  │
│                                                                             │
│  ③ Injected "environment memory"                                            │
│     CLAUDE.md, user preferences, memory files, date, working directory      │
│                                                       ← re-injected every   │
│                                                         request             │
│                                                                             │
│  ④ Conversation history and tool results (the growing body)                 │
│     user/assistant messages + every tool call and its return                │
│     ├─ File contents, grep results, shell output (often thousands of        │
│     │  tokens at a time)                                                    │
│     ├─ Error messages and stack traces (deleting them is a contested call)  │
│     └─ The model's own earlier reasoning and plans (externalized artifacts  │
│        like todo lists)                                                     │
│                                                                             │
│  ⑤ This turn's user input                                                   │
└─────────────────────────────────────────────────────────────────────────────┘

The most important insight in this diagram: except for slot ⑤, the harness decides the content of every slot. How the system prompt is written, which tools are exposed, which fragments of memory files get injected, whether tool results flow back verbatim or get trimmed, when history gets compressed — the model has no say in any of it. That's what "context engineering is the harness's core responsibility" means.

The converse explains why the same model behaves completely differently in different harnesses: the model is identical, but what it sees at each step is not.

Four ways context fails ​

Drew Breunig's How Long Contexts Fail sorts long-context failures into four categories, and this taxonomy has become the de facto vocabulary for discussing the problem. One by one, with the original evidence:

Context poisoning: an error enters the context and keeps getting cited ​

Once a hallucination or mistake is written into the context, later generations treat it as fact and cite it repeatedly — a self-reinforcing loop. Google DeepMind's Gemini 2.5 technical report documented a classic case: their Pokémon-playing Gemini agent would occasionally hallucinate game state, and once a wrong state entered the "goals/summary" region, the agent would "persist in working toward impossible or irrelevant goals" and "often took a long time to recover."

The lesson for harnesses: every piece of information entering the context needs credibility management. Garbage returned by tools, wrong assumptions the model itself wrote earlier, unreliable conclusions handed back by subagents — all potential poison sources.

Context distraction: the context drowns out what training taught ​

Once the context gets long enough, the model leans heavily on what's in front of it and underuses what it learned in training. The Gemini 2.5 report observed that once context significantly exceeded 100k tokens, agents tended to repeat actions that had already appeared in history rather than synthesize a new plan. Databricks' long-context research (cited by Breunig) puts harder numbers on it: Llama 3.1 405B's accuracy starts degrading around 32k tokens, and smaller models degrade even sooner.

This is the strongest counter-evidence to the "bigger window is always better" myth: models start getting dumber long before the window is full. Where very long windows genuinely shine is summarization and factual retrieval — not long-chain generative reasoning.

Context confusion: irrelevant content contaminates the response ​

Whatever you put into the context, the model has to "process" it — irrelevant documents and unused tool definitions all shape the output. The most empirical evidence here comes from tool counts: the Berkeley Function-Calling Leaderboard shows that every model performs worse when given more than one tool, and models will occasionally force a tool call even when no relevant tool exists in the scenario. Breunig also cites a GeoEngine benchmark experiment: a quantized Llama 3.1 8B failed the task when given 46 tool definitions (nowhere near the context limit), and succeeded once that was cut to 19.

A sobering datapoint for the MCP gold rush: stuffing hundreds of tool descriptions into the context doesn't empower your agent — more often it manufactures confusion. This is exactly why the tool system keeps insisting on "few, sharp, clearly bounded."

Context clash: the context contradicts itself ​

Worse than "irrelevant" is "contradictory": incoming information directly conflicts with what's already in the context. A Microsoft–Salesforce paper, LLMs Get Lost in Multi-Turn Conversation (arXiv:2505.06120), delivered a startling quantified result: change the exact same task information from "all at once" to "sharded across multiple turns," and every top model's performance drops by 39% on average — o3 falls from 98.1 to 64.1. The authors' diagnosis: models make assumptions in early turns, commit prematurely to incomplete solutions, and those wrong answers stay in the context, continuously shaping later generation — "once a model takes a wrong turn, it gets lost and doesn't find its way back."

This is doubly bad news for agents

An agent's context is inherently "assembled from fragments": information from different documents, different tools, and different subagents arrives over time, so the odds of contradiction are far higher than in a single-turn setting. And the agent's own failed attempts in early turns are a clash source too. Multi-turn degradation isn't a chatbot-only disease — it's baked into how agents work.

Here's the four modes as a quick-reference table for diagnosing anomalous agent behavior:

Failure modeOne-line definitionTypical symptomsFirst-line fix
PoisoningA hallucination/error enters the context and keeps getting citedObsessed with impossible goals; errors self-reinforceCredibility management for sources; verify key facts when summarizing
DistractionContext too long; the model repeats past actionsStops generating new plans; spins in placeCompaction; prune redundant history
ConfusionIrrelevant content gets processed as if it matteredCalls irrelevant tools; cites irrelevant documentsTrim the tool set (loadout); retrieve only what's needed
ClashInformation inside the context contradicts itselfAnchored to early wrong answers; can't self-correctQuarantine conflicting sources; divide and conquer with subagents

Note that they frequently co-occur: a poisoning event creates clashes in later turns, and an overlong history feeds both distraction and confusion. The value of diagnosis is picking the right lever — a bigger window doesn't help with any of the four.

Compression and truncation: survival strategies for long tasks ​

Run a task long enough and the context approaches its ceiling. The harness has three moves: compaction, pruning/masking, and offloading.

Compaction (summarize and restart). When the context nears the window limit, hand the entire conversation history to the model for summarization and start a fresh window from the summary. Claude Code's implementation is the industry reference point: the model keeps architectural decisions, unresolved bugs, and implementation details, discards redundant tool output, then keeps working with the compressed context plus the five most recently touched files. Anthropic's tuning advice is refreshingly practical: first maximize recall on complex trajectories (make sure the summary captures every relevant fact), then iterate on precision (cut the excess) — the things over-aggressive compression throws away, you only discover mattered ten steps later.

Tool result clearing / masking. The lightest, safest form of compression: once a tool call is buried deep in history, the model rarely needs its raw output anymore — replace old tool results with placeholders. Anthropic has shipped this as a platform feature on Claude. Note that this is "clearing," not "deleting": the call record (what was called, with what arguments) is kept; only the bulky return body is hidden.

Restorable offloading. In their build retrospective, Manus proposed a more radical principle: any irreversible compression is risky, because you can't know which observation becomes critical ten steps later. Their solution: "treat the file system as the ultimate context" — web content can leave the context as long as the URL remains; document content can be omitted as long as the path still exists in the sandbox. Compression becomes restorable: what you discard is content, what you keep is the index. This is the same philosophy as the note-style externalization in memory systems (NOTES.md, todo files): the context holds only "what's needed now"; everything else lives outside as retrievable references.

One easily-missed hedge, also from Manus: don't clear errors out of the context. Erasing traces and hiding failures robs the model of the evidence that "this was tried, and it didn't work" — so it steps on the same rake again. Failed tool calls, along with their errors, staying in the context is a precondition for self-correction. What a compression strategy should discard is "redundant detail from successes," not "the facts of failures."

KV cache: why a stable prefix is money ​

Context engineering isn't just a quality game; it's a cost game. Agent workloads have a distinctive economics: extremely long inputs, extremely short outputs. Manus reports an average input-to-output ratio of roughly 100:1 — every step re-prefills the entire history just to produce a tool call a few dozen tokens long.

This makes the KV cache hit rate one of the single most important metrics for a production agent (Manus's own words). Cache hits directly determine TTFT and cost: on Claude Sonnet, cached input tokens run about $0.30 per million versus about $3 uncached — a 10× spread. And the mechanics of the KV cache impose one hard constraint: the prefix invalidates from the first differing token onward.

From that, three engineering disciplines follow:

  1. Keep the prefix absolutely stable. The classic self-inflicted wound is putting a timestamp accurate to the second at the top of the system prompt — it changes every step, killing the entire cache. Volatile information like dates goes at the tail, or nowhere.
  2. Append-only context. Never go back and rewrite past actions or observations; serialization must be deterministic (even unstable JSON key order can silently bust the cache).
  3. Tool definitions live at the head → the tool set can't change mid-run. Most models serialize tool definitions at the front of the context; dynamically adding or removing tools invalidates the cache for everything after them. Manus's workaround is "mask, don't remove": leave tool definitions in place and restrict the selectable action set at decode time with a logits mask — you get "only these tools are available right now" while keeping the cache intact.

How this constraint shapes product design

The "stable prefix" explains why mainstream coding agents ship a static, long-form system prompt with dynamic content appended at the tail, rather than assembling the prompt fresh each step. It also explains why the todo list in planning lives in the conversation flow (appendable, overwritable) instead of being baked into the system prompt (which would bust the cache). For every block of content in the context, its position is itself an engineering decision.

Recitation: a cheap way to steer attention ​

Position within the context is not neutral. The "lost in the middle" phenomenon tells us models use what's at the beginning and end of the context best, and the middle worst. That gives the harness a lever that moves content, not architecture — recite the most important things to the tail of the context.

Manus turned this into an explicit mechanism: on complex tasks, the agent creates a todo.md and rewrites it after every step. A Manus task averages around 50 tool calls, so the original goal can easily drown under dozens of observations; continuously rewriting the todo list amounts to reciting the global plan to the most recent position in the context, pushing it back into the model's short-range attention window and counteracting goal drift. Note there is no model-side trick here at all — it's purely biasing attention by controlling "what sits at the tail of the context right now."

This explains, from a context-engineering angle, why TodoWrite works in planning and task decomposition: the todo list isn't a progress bar for the user — it's an artifact deliberately placed in a high-attention zone. The same principle underlies "keep recently touched files after compaction" and "summarize, then open a new window" — deciding what appears at the end of the context is deciding what the model will "think of" next.

Retrieve only what's needed: RAG vs. just-in-time ​

There are two basic routes for getting relevant information into the context:

Pre-retrieval (RAG-style): before inference, use embeddings to pull out plausibly relevant documents and inject them all at once. Fast and deterministic; the downside is that it's a guess — retrieval happens before the model starts understanding the task, and a wrong guess means irrelevant content eating the attention budget (the confusion problem above).

Just-in-time retrieval: keep only lightweight identifiers in the context (file paths, URLs, queries) and let the agent load what it needs at runtime through tools. Anthropic notes the industry is clearly shifting toward this route — Claude itself uses it for complex analysis over large databases: the model writes targeted queries, stores results, and pulls only the slices it needs with commands like head/tail, never moving the full data object into the context.

Just-in-time retrieval also pays a hidden dividend: progressive disclosure. File size hints at complexity, naming conventions hint at purpose, timestamps hint at freshness — tests/test_utils.py and the same-named file under src/core_logic/ mean entirely different things. The agent builds understanding layer by layer off the file system as an "external index," keeping only the necessary subset in working memory.

The cost is just as clear: runtime exploration is slower than precomputed retrieval, and if the tools and heuristics are badly designed, the agent burns context on dead ends. So the realistic answer is hybrid, and Claude Code is the exemplar — stable information like CLAUDE.md is injected up front, while code content is fetched just-in-time via glob/grep, which conveniently sidesteps stale indexes and complex syntax trees.

DimensionPre-retrieval (RAG)Just-in-time
TimingInjected once, before inferenceLoaded on demand at runtime
LatencyLow (precomputed)High (multi-turn exploration)
RelevanceDepends on the retriever's "guess"The model judges for itself — usually more accurate
Context costAll at once; may include irrelevant materialGrows on demand, but exploration itself costs tokens
Data freshnessOnly as fresh as the indexAlways reads current state
Best forRelatively static content (legal, financial documents)Dynamic environments (codebases, databases, the web)

The same logic applies to the tool set itself: in the follow-up How to Fix Your Context, Breunig calls it tool loadout — like a video-game loadout screen, give only the tool definitions relevant to the current task. The RAG-MCP experiment found that past ~30 tools, descriptions start overlapping and selection accuracy collapses; retrieval-based tool selection improved selection accuracy by roughly 3×. But note the tension with KV cache discipline: the loadout should be fixed at the start of a session/task, not shuffled mid-loop.

Design takeaways ​

The whole page, condensed into an executable checklist:

  • Manage context like a budget. Every token spends attention; aim for "the smallest high-signal set," not "fill the window."
  • Failure shows up before the window fills. Diagnose anomalous behavior with the four failure modes (poisoning / distraction / confusion / clash) instead of reflexively reaching for a bigger window.
  • Compress in layers: clear old tool results first (lossless), compaction only if that's not enough (lossy), offload to file/URL indexes before summarizing away (restorable).
  • A stable prefix is cost discipline. System prompt and tool definitions frozen at the head; dynamic content appended at the tail; never a timestamp at the top.
  • Just-in-time by default; pre-inject only explicit exceptions. Stable, always-needed material (specs, memory files) goes up front; everything else gets an identifier, and the model pulls.
  • Keep errors in the context. Failed trajectories are the evidence a model needs to self-correct; when compressing, preserve "what didn't work" first.
  • Quarantine is the strongest lever. When one context can't be managed no matter what, split the task across subagents with clean windows — Anthropic's multi-agent research system works precisely by having each subagent explore tens of thousands of tokens and return only a distilled thousand or two.

Further reading ​

References ​