Skip to content

Context Engineering

At a glance The paradigm shift from prompt engineering to context engineering — why context is a scarce resource, how the four strategies of writing/selection/compression/isolation land in practice, and the production-proven techniques for session compaction, tool masking, and KV-cache optimization validated by Manus and Claude Code.

Context Engineering ​

In June 2025, Andrej Karpathy posted on X arguing that "context engineering" should replace "prompt engineering" as the name for the real work of building LLM applications. Tobi Lutke (Shopify's CEO), Simon Willison, and others quickly piled on, and within a month the term was industry consensus. By the second half of 2025, Anthropic, LangChain, and Manus had each published long-form write-ups of their production context-engineering practice — and all three independently reached nearly the same conclusion: most agent failures are context failures, not model failures.

This page nails down three things: why context is a scarce resource that must be budgeted carefully; the four management strategies the industry has converged on; and the battle-tested techniques you can only learn by getting burned in production.

1. Definitions: from Prompt Engineering to Context Engineering ​

The paradigm shift ​

Prompt engineering answers: "how do I write this instruction for the best effect?" It carries a hidden assumption — the task is one-shot and the input is controllable, and what you're polishing is that one system prompt.

Agents break that assumption. An agent runs tens to hundreds of turns inside the Agent Loop, and what the model sees changes every turn: system prompt, tool definitions, tool results, retrieved documents, new user messages, the accumulated conversation. Cognition (the Devin team) put it most bluntly: context engineering is "the number one job of the engineer building an AI agent."

Anthropic's engineering post Effective Context Engineering for AI Agents (September 2025) offers a widely cited definition:

Context engineering is the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference — including everything beyond the prompt that can land in the context.

Karpathy's framing is more viral: he likens the LLM to a new kind of operating system — the LLM is the CPU, the context window is RAM. Just as an OS carefully schedules what enters memory, context engineering is "the delicate art and science of filling the context window with just the right information at each step."

One picture of what's actually inside the context ​

┌─────────────────── Context Window (finite, scarce) ──────────────────┐
│                                                                      │
│  ┌──────────────┐  ┌──────────────┐  ┌─────────────────────────┐    │
│  │ System Prompt│  │  Tool defs   │  │  Few-shot examples      │    │
│  │ (role/rules) │  │  (schemas)   │  │  (canonical examples)   │    │
│  └──────────────┘  └──────────────┘  └─────────────────────────┘    │
│                                                                      │
│  ┌─────────────────────────────────────────────────────────────┐    │
│  │ Message history (user / assistant / tool_result, accruing)  │    │
│  │   — the only part that only ever grows, and the part that   │    │
│  │     most needs managing                                     │    │
│  └─────────────────────────────────────────────────────────────┘    │
│                                                                      │
│  ┌──────────────┐  ┌──────────────┐  ┌─────────────────────────┐    │
│  │ RAG results  │  │ Memory       │  │ Sub-agent summaries     │    │
│  │              │  │ injections   │  │                         │    │
│  └──────────────┘  └──────────────┘  └─────────────────────────┘    │
│                                                                      │
└──────────────────────────────────────────────────────────────────────┘
  Before every Agent Loop iteration, re-decide: which of these tokens
  should enter, stay, be compressed, or be dropped

LangChain's Context Engineering for Agents groups the context to be managed into three classes: instructions (prompts, examples, tool descriptions), knowledge (facts, memories), and tool feedback (tool-call returns). Each class has a different lifecycle and update frequency, and each calls for different management tactics.

Its relationship to prompt engineering

Context engineering doesn't replace prompt engineering; it's a superset. The system prompt still needs to be written well (see Prompt Engineering) — it's just a small fraction of the total work. A quick test: if your debugging method is endlessly rewriting the prompt and never inspecting "which tokens did the model actually see at step 14," you're still doing agent work with a prompt-engineering mindset.

2. The Economics of the Context Window: Attention Is Scarce ​

Why not just wait for bigger windows and stuff everything in? Because evidence on three levels says: the longer the context, the worse the model performs — and the degradation is continuous, starting from the very first token.

The attention budget and context rot ​

Transformer attention relates n tokens pairwise, producing n² relationships. As context grows, the model's capacity to capture those relationships gets spread thin — Anthropic calls this the finite attention budget, and every new token spends it. Add one more fact: short sequences vastly outnumber long ones in training data, so models are naturally under-experienced with very long contexts. The result is a gradient-like slide in performance, not a cliff at the window limit.

Chroma Research's 2025 technical report Context Rot turned the phenomenon into hard numbers: testing 18 models (including GPT-4.1, the Claude 4 family, Gemini 2.5, Qwen3), holding task difficulty constant and increasing only input length, they found:

  • Even on a near-trivial task like "copy a string of repeated words," every model degraded as input grew — and the degradation was non-uniform and hard to predict;
  • A single distractor lowered accuracy, and stacking four made it worse — in long contexts, irrelevant-but-plausible content is toxic;
  • On the LongMemEval conversational-memory benchmark, models given "focused ~300 tokens of relevant context" performed significantly better than with "the full ~113k-token raw history" — irrelevant context forces the model to retrieve before reasoning, and the retrieval step itself costs points.

"Lost in the middle": position charges rent too ​

The earlier foundational study is Liu et al.'s 2023 paper Lost in the Middle: How Language Models Use Long Contexts (arXiv:2307.03172, later in TACL 2024). Core finding: models use information at the beginning and end of context best; recall for the middle drops noticeably, tracing a U-shaped curve. When the key document sits in the middle of a long context, some models perform worse than a closed-book baseline with no document at all.

Direct engineering corollaries:

  • Put the most important instructions at the start and the most critical recent information at the end — the middle "waist" is attention's lowland;
  • Don't dump 20 retrieved documents in at once — ordering, truncation, and compression matter more than recall volume (see RAG);
  • "Throw it all in; the model will find it" does not survive long contexts.

The four failure modes of long context ​

Drew Breunig's How Long Contexts Fail (June 2025) gives a failure taxonomy the industry has widely adopted — LangChain's official summary cites it too:

Failure modeMeaningTypical scenario
Context PoisoningA hallucination or error enters the context and gets cited by later stepsThe agent misjudges a file's structure; every later decision builds on the wrong belief
Context DistractionThe context is so long the model over-attends to history and under-uses what it learned in trainingAfter a wrong fix appears several times in history, the model stops trying new paths
Context ConfusionIrrelevant content contaminates the outputWith an oversized toolset, the model calls a similarly named tool for the wrong purpose
Context ClashInformation inside the context contradicts itselfAn early assumption conflicts with a later tool result, and the model wavers between them

Note the counterintuitive corollary: much of the felt "the model got dumber" is context-management debt compounding. The model didn't change — the tokens it sees each step got dirtier. Which means most of these problems can be fixed by engineering, without upgrading the model. That's exactly why context engineering exists.

A big window is not a get-out-of-jail card

Models like Gemini offer million-token windows, eliminating "it won't fit" — but not attention dilution, context rot, or cost, which grow roughly linearly with context length. A million-token window is a capacity upgrade, not a quality exemption. With poor context quality, a bigger window just makes your mistakes more expensive.

3. The Four Strategies: Write, Select, Compress, Isolate ​

Harrison Chase's team at LangChain published Context Engineering for Agents in mid-2025 with the most widely used classification framework. It asks one question from the perspective of information flow: relative to the context window, is each token being written out, selected in, compressed, or isolated elsewhere? The answers map to four strategies:

StrategyOne-line definitionTypical tactics
WriteStore information outside the context window; fetch it when neededScratchpads, file notes, cross-session memory
SelectPull only the currently needed tokens in from external sourcesRAG, on-demand tool loading, memory retrieval
CompressKeep the tokens the task needs; discard the restSession summaries, history trimming, tool-result clearing
IsolateSplit context across different agents/threads so none pollutes the othersSub-agents, multi-threaded state, environment isolation

Write: put it out there ​

Things inside the window get washed away; things outside don't. An agent taking notes as it works (a scratchpad) directly imitates how humans work: Claude Code maintains a todo list; Anthropic's multi-agent research system has the LeadResearcher write its plan into Memory so direction survives a 200k-token truncation.

Implementation can be lightweight: a single write-file tool call, or a persistent field on a runtime state object. The point isn't the carrier — it's explicitly externalizing anything that must survive compaction rather than hoping it stays in the message history. Writing memories across sessions is a separate topic; see Memory Systems.

Select: bring the right things in ​

Select answers "what does this step need?" Two directions:

  • Knowledge selection: RAG retrieves relevant chunks instead of dumping everything; memories are recalled by relevance;
  • Tool selection: don't hang 50 tools on the agent — load on demand. Tool definitions themselves cost tokens and manufacture Context Confusion. Anthropic's judgment is blunt: if a human engineer can't say which tool fits a scenario, an agent won't do better.

An increasingly mainstream stance is just-in-time context: don't pre-retrieve all data; give the agent lightweight references (file paths, stored queries, URLs) and let it explore and pull data into context progressively at runtime through tools. Claude Code is a hybrid: CLAUDE.md is injected at startup, everything else is navigated on demand with glob/grep — avoiding stale indexes while leaving the exploration decisions to the model.

Compress: make it smaller ​

Details in the next section's practice part. The core moves: summarization, trimming, tool-result clearing. Compression is the first lever on long tasks — and the easiest to overdo, because aggressive compression discards information that looked unimportant then but turns critical later.

Isolate: keep them apart ​

Keep a main agent's context a clean, high-level plan, and subcontract the dirty work (massive searches, reading dozens of files, running commands) to sub-agents. Each sub-agent explores tens of thousands of tokens in its own clean window and returns only a 1,000–2,000-token distilled summary. Anthropic's multi-agent research system beats a single agent substantially on complex research tasks precisely through this pattern. The architectural expansion is in Multi-Agent Systems.

The four strategies are not mutually exclusive

Production systems almost always mix all four. Claude Code: writes (todo list), selects (just-in-time file navigation), compresses (compaction), isolates (Task sub-agents) — the works. The framework's value is building your classification intuition: when a context problem appears, categorize it first, then pick the lever.

4. Battle-Tested Techniques ​

Everything in this section comes from public production practice at frontline teams, not blog speculation. The sources cluster in three places: Anthropic's engineering blog, Manus's Context Engineering for AI Agents: Lessons from Building Manus (July 2025), and LangChain's how_to_fix_your_context example repo.

Session compaction ​

When a long session nears the window limit, hand the message history to the model for summarization and start a new window from the summary. Claude Code's implementation: keep architecture decisions, unresolved bugs, and implementation details; drop redundant tool output; continue with the compacted context plus the five most recently touched files.

Anthropic's tuning methodology is worth copying outright: maximize recall first, then iterate toward precision. Make sure the compaction prompt captures every possibly relevant item in the trace (better to over-retain), then progressively delete categories proven redundant. The safest low-hanging fruit is clearing old tool calls and results deep in the history — those tools already ran, and the model doesn't need to see the raw output again. Anthropic later productized this as the platform-level tool result clearing feature.

A minimal working compaction loop:

python
# Framework-agnostic compaction skeleton; pair it with your framework's state management in real projects
def should_compact(messages, token_counter, threshold=0.8, window=200_000):
    """Trigger compaction when used tokens exceed 80% of the window"""
    return token_counter(messages) > threshold * window


def compact(messages, llm, keep_recent=6):
    """
    Compaction strategy:
    1. Keep the most recent `keep_recent` messages verbatim (preserves conversational flow)
    2. Hand everything older to the model for a high-recall summary
    3. Summary + recent messages = the new context
    """
    head, tail = messages[:-keep_recent], messages[-keep_recent:]

    summary = llm.complete(
        "You are a conversation compactor. Summarize the agent work history below. You must keep: "
        "(1) confirmed architecture/design decisions; (2) unresolved problems and error messages; "
        "(3) key file paths, variable names, and interface contracts; "
        "(4) preferences and constraints the user stated explicitly. "
        "You may discard: raw output of already-completed tool calls, repeated exploration.\n\n"
        f"Message history:\n{format_messages(head)}"
    )

    # Carry the summary in a system message, followed by the uncompressed recent messages
    return [{"role": "system", "content": f"[Summary of earlier conversation]\n{summary}"}, *tail]

The point isn't the code; it's the trade-offs: the trigger threshold, how much raw text to retain, and which categories the summary prompt names explicitly — these three knobs decide whether compaction is a lossless extension of life or gradual amnesia. Tune the summary template against real failure traces; don't write it from imagination.

Tool-result masking, and "don't dynamically add/remove tools" ​

Manus shared a counterintuitive lesson: do not add or remove tool definitions mid-run. Two reasons: first, tool definitions usually sit at the front of the context, so changing them invalidates the KV cache for every prior turn (next item); second, the message history now references tools that no longer exist, confusing the model.

Their alternative is masking: keep the full toolset stable and control the currently available actions through constrained decoding (masking unavailable tool tokens at the logits level). The context doesn't change, the cache stays valid, but the model's actual choice space narrows precisely. Similarly, bulky tool results deep in the history (say, an entire ingested log file) can be replaced with placeholders once confirmed unnecessary — context length unchanged, information density restored.

The filesystem as external context ​

Manus and Anthropic independently converged on the same pattern: treat the filesystem as an external context of unbounded capacity. The agent doesn't stuff whole datasets into the window; it holds references — file paths, stored queries, URLs — and loads on demand with tools. Claude Code can analyze databases and logs larger than the window with head/tail, never reading the full data object into context.

This pattern has an underappreciated dividend: the filesystem's metadata is itself context. tests/test_utils.py and src/core_logic/test_utils.py share a name but not a meaning; directory hierarchy, naming conventions, and timestamps all help the agent judge relevance — these signals are free, so use them. Manus's post even offers a quantifiable mental baseline: they prefer having the model offload recoverable, bulky content to the filesystem, because "as long as the URL is there, web content can be dropped from the context."

KV-cache hit-rate optimization ​

The most engineering-dense section of Manus's post. The KV cache lets tokens sharing a prefix skip recomputation; a cache hit versus a miss can differ by an order of magnitude in inference cost (Manus reported roughly a 10× price gap between cached and uncached input tokens for Claude Sonnet at the time). Every agent turn is "prefix + delta," so cache hit rate is the lifeline of cost.

Three practical principles:

  1. Keep the prefix stable. Any single-byte change to the system prompt or tool definitions invalidates every cache entry behind it. Don't even put timestamps in the prefix — a timestamp precise to the second at the top of the system prompt is a self-inflicted cache wipe.
  2. Append only; never rewrite. The message history is append-only. Make sure serialization is deterministic: the same JSON content with a different key order silently invalidates the cache too.
  3. Mark cache breakpoints manually. Some platforms require explicit cache breakpoints; place them with prefix reuse in mind.

Think before you compress

The KV-cache view gives "when to compress" a sharper criterion: compaction is a wholesale prefix rewrite that invalidates all old cache. So compress late (don't touch it while performance and cost still have headroom), but when you do compress, compress thoroughly — frequent small compactions are the worst strategy, paying the full-prefix price every time.

Recitation ​

Another cheap trick from Manus: have the agent recite the todo list or current plan into the tail of the context. Because models naturally attend more to the end of the context (recall Lost in the Middle's U-shaped curve), periodically "restating" the global goal at the most recent position markedly reduces goal drift on long tasks. Claude Code's todo-list tool serves exactly this role — it's not just a project-management tool; it's an attention-manipulation tool.

A cheat sheet for long-horizon tasks: pick one of three ​

Anthropic draws clear applicability boundaries for three long-horizon techniques, directly usable as selection criteria:

  • Compaction fits tasks needing lots of back-and-forth interaction where conversational continuity comes first — it preserves "where the conversation stands";
  • Structured notes fit iterative development with clear milestones — they preserve "which step we're on";
  • Sub-agent isolation fits complex research and analysis that can be explored in parallel — it preserves "the main line from being drowned by side quests."

Real systems usually chain them: sub-agents each take their own notes, returned summaries enter the main context, and the main context compacts when it nears capacity. Understanding what each technique preserves matters more than memorizing the techniques.

5. Anti-Patterns ​

These patterns recur constantly in production traces; each carries a real cost:

1. Context bloat. Stuffing in everything "that might help": full chat logs, whole files, dozens of pages of retrieval results. The cost is threefold — linearly rising spend, diluted attention, accelerated context rot. The test is simple: if deleting a piece of context doesn't change output quality, it never belonged there.

2. Instruction accretion. Every user complaint adds a rule to the system prompt; six months later it's an internally contradictory patchwork legal code — "be detailed" coexists with "be concise," "proactively clarify" fights "don't ask too much." Anthropic's advice: replace exhaustive rule lists with canonical examples (carefully chosen, diverse exemplars) — examples say in one shot what rule-stacking can only mangle, and rule piles manufacture Context Clashes. Aim for the system prompt to be "the minimal information set that fully defines behavior" — minimal, not necessarily shortest.

3. Context poisoning. One hallucination, one wrong tool return, one stale retrieval result enters history and gets cited as fact for dozens of turns. Defenses: validate key facts before writing them to the scratchpad; when a wrong belief is found in history, correct it explicitly in the context (don't silently hope the model notices the contradiction); keep source attribution on external retrieved content.

4. Few-shot mode lock-in. Manus calls it "don't get few-shotted": LLMs are superb mimics, and if the context fills with similar action-observation pairs, the model will mindlessly continue the pattern even when it's no longer optimal — reviewing 20 résumés in batch, by the 15th the output format and approach are locked by the previous 14, and per-case judgment is gone. The fix isn't more examples; it's introducing variation: vary the phrasing structure, rotate the order, control the density of same-type samples.

5. Error erasure. Intuition says delete failed attempts from history to keep the context "clean." Manus's production experience is the opposite: keeping errors and recoveries is one of the agent's most effective learning signals. Only by seeing the full trajectory of "this command failed → that approach worked" will the model avoid the same trap in similar situations. An agent whose errors are erased will step on the same rake repeatedly.

Don't over-engineer your context

The opposite failure deserves equal caution. Anthropic's advice is "build the simplest thing that works": short sessions, few tools, and single-turn tasks don't need scratchpads, sub-agents, or compaction pipelines. The four strategies are a toolbox, not a checklist. Building a compaction system for tasks under 10 turns is itself a case of engineering-resources context bloat.

6. The Boundary with Memory Systems ​

Context engineering and memory are constantly conflated; the boundary is worth drawing:

  • Context engineering governs "this inference": what goes in the window, how, and how much. Timescale: a single task, a single session.
  • Memory governs "across sessions": which information deserves to outlive the task, and how to recall it at next startup. Timescale: days, months, the user's lifetime.

The two systems meet at the "write" and "select" actions: which notes in the scratchpad get promoted to long-term memory when the session ends? And of everything stored in long-term memory, which few entries should this step recall into the window? The former is a memory system's write-policy question; the latter is context engineering's selection question.

Put differently: the memory system decides what's in the warehouse; context engineering decides what's on the workbench. However big the warehouse, the wrong items on the workbench still mean bad work; however well-run the workbench, an empty warehouse accumulates nothing across tasks. Design the two separately and tune them together — the deep dive is in Memory Systems.

References ​