Skip to content

Planning & Task Decomposition

At a glance A deep dive into how AI agents plan — a comparison of three modes (no planning, plan-and-execute, interleaved planning), how the todo list works as explicit working memory, and whether explicit planning still matters in the era of reasoning models.

Planning & Task Decomposition ​

Planning answers this question: before (or while) executing, should the model write down "what to do next" explicitly — and in what form?

In the harness context, planning isn't some abstract cognitive capability of the model. It's a concrete set of system design decisions: where the plan lives (text in the context, a structured tool call, or a fixed flow in code), who generates it (the model, another model, or a human), how often it gets updated, and what happens when it goes stale. These decisions directly determine an agent's performance, cost, and controllability on long tasks.

Do agents need explicit planning? ​

This is one of the most contested questions in agent engineering over the past three years, and both camps have heavyweight proponents:

Against: planning is redundant scaffolding. A purely reactive agent — receive an observation, directly decide the next action — has the simplest structure, and every step is decided against the freshest environment state, so it never gets held hostage by a pre-written plan that may already be obsolete. The painful experience of the AutoGPT era (released March 2023) backs this camp up: the model would solemnly draft a grand plan, then grind for half an hour on a single failed step — the plan itself became a sunk cost.

For: without an explicit plan, long tasks inevitably drift. An LLM's attention and working memory both live inside the context window; after a few dozen tool calls, "what was the original goal?" gets buried under a pile of tool output. The essence of an explicit plan is externalizing the goal — writing it into a checklist the model can re-read at any time to fight context drift. Claude Code says it outright in its system prompt: skipping TodoWrite when planning tasks "may forget to do important tasks — this is unacceptable" (per a reverse-engineering analysis of its prompts).

A footnote to this debate

Curiously, Claude Code's official documentation only mentions that the TodoWrite tool exists — it never documents the tool's core "proactive planning" behavior, to the point that a user filed issue #6968 complaining about the gap between docs and behavior. Which says something indirectly: in these products, explicit planning isn't an optional "feature" — it's part of how the agent fundamentally operates.

The position of this page (which is also where mainstream products have landed in practice): the debate shouldn't be "whether to plan" but "at what granularity, in what form, and at what moments the plan gets written and revised." The rest of this page lays out that design space with three modes.

Three planning modes ​

Mode 1: No planning (purely reactive) ​

The most primitive agent loop: observe → act → observe → act… There's no separate planning phase; every step is the model directly picking an action based on the current context.

text
┌─────────────────────────────────────────────┐
│           Purely reactive (ReAct-style)     │
│                                             │
│   Observe ──> think ──> act ──> observe ──> …│
│                                             │
│   The plan exists only in the model's       │
│   "thinking right now" — no persisted       │
│   plan state whatsoever                     │
└─────────────────────────────────────────────┘

The exemplar is ReAct (Reasoning + Acting; paper by Yao et al., 2022, accepted to ICLR 2023): the model alternates between emitting a reasoning trace and an action. The thinking is the planning — but it's ephemeral, discarded as soon as it's generated.

  • Pros: minimal structure; every decision is based on the latest state, so it adapts to environmental change naturally; no "stale plan" problem.
  • Cons: no global view. The model easily falls into local loops (retrying the same failing action), and on long tasks it gradually forgets the original goal. Thinking and acting are coupled, so you can't separately review "what it intends to do."

The base loops of SWE-agent and OpenHands are essentially this mode, heavily reinforced with tools and prompt engineering.

Mode 2: Plan-and-Execute ​

Planning and execution are split into two modules and two LLM calls: a Planner produces a complete multi-step plan up front, an Executor carries it out step by step, and an optional Replanner intervenes on failure.

text
┌──────────────────────────────────────────────────┐
│              Plan-and-Execute                    │
│                                                  │
│  Task ──> ┌─────────┐  full plan  ┌──────────┐   │
│           │ Planner │ ─────────> │ Executor │   │
│           │(big LLM)│  step 1..N │(small OK)│   │
│           └─────────┘            └────┬─────┘   │
│                 ▲                     │         │
│                 │ replan on           │         │
│                 │ failure/drift       │         │
│           ┌─────────┐                 ▼         │
│           │Replanner│ <───────  results        │
│           └─────────┘                           │
└──────────────────────────────────────────────────┘

The lineage traces back to the Plan-and-Solve prompting method (paper by Wang et al., 2023 — have the model draft a plan, then solve step by step) and BabyAGI (a task queue + task-creating-agent loop architecture). LangChain's official blog (February 2024) systematically summarized how to implement plan-and-execute agents.

  • Pros: the whole plan is visible and auditable before execution begins (a human can approve it); the planning phase isn't re-invoked on every action, which saves tokens — engineering teams have reported 30–60% token savings on multi-tool tasks (reference — anecdotal experience, not a rigorous benchmark); it enables model tiering: an expensive big model plans, a cheap small model executes.
  • Cons: plan brittleness is the fatal flaw. The plan is generated at T0 from incomplete information; by step 3 the world has changed, but the plan doesn't know. Triggering replanning frequently enough to compensate eats the cost savings. The more detailed the plan, the faster it shatters.

Mode 3: Interleaved planning ​

Planning isn't a one-time event but a continuous activity spanning the whole session: the agent maintains a persistent plan artifact (usually a todo list) and keeps editing it during execution — mark a step done, insert a new task when new information shows up, cross out and rewrite when the direction is wrong.

text
┌──────────────────────────────────────────────┐
│          Interleaved (Claude Code-style)     │
│                                              │
│   ┌───────────────┐                          │
│   │   Todo List   │ <── status written back  │
│   │(explicit      │     after every action   │
│   │ working memory)│                         │
│   └──────┬────────┘                          │
│          │ visible before every decision     │
│          ▼                                   │
│   Observe ──> think ──> act ──> observe ──> … │
│                                              │
│   Planning and execution share one agent     │
│   loop; the plan is "alive" and evolves      │
│   with the environment                       │
└──────────────────────────────────────────────┘

Claude Code's TodoWrite + execution loop is the benchmark implementation of this mode; its plan mode (draft a plan, submit it for human approval via the ExitPlanMode tool, then start work) layers a "human approval at critical junctures" gate on top of interleaving.

  • Pros: global view and adaptability at once — the plan is always in the context (fighting goal drift) yet can be revised at any moment (fighting plan brittleness); users see progress in real time, and the cost of intervention is low.
  • Cons: an extra layer of tool-call overhead (every status update is a tool call); it relies on the model diligently maintaining the list, and the model may forget to update status (more below); implementation is more complex than pure reaction.

The three modes compared ​

DimensionPurely reactivePlan-and-ExecuteInterleaved
Plan representationEphemeral reasoning traceComplete plan generated oncePersistent, continuously revised task list
When the plan is madeImplicitly, every stepOnce, before executionDynamically, throughout the session
Global viewWeakStrong (but goes stale)Strong (and stays fresh)
Environmental adaptabilityStrongestWeak (depends on replanner)Strong
Auditability / approvabilityPoorBest (fully visible up front)Good (visible in real time)
Token costLowLow (if the planner isn't re-invoked)Medium-high (list rides along every step)
Typical examplesReAct, base SWE-agent loopBabyAGI, LangChain plan-and-executeClaude Code (TodoWrite)

A decision framework

The shorter the task and the more dynamic the environment, the better reactive fits; the more certain the steps and the more you need up-front approval, the better plan-and-execute fits; long task + uncertain environment + humans need to jump in anytime — interleaved is the best engineering compromise today, which is why mainstream coding agents have all converged on it.

The todo list: explicit working memory as a mechanism ​

The todo list deserves its own dissection because it looks trivial but is actually the most elegant piece of design in the interleaved mode: it reduces a "planning problem" to a "tool-call problem."

How Claude Code's TodoWrite actually works ​

Based on reverse-engineering of Claude Code's prompts and tool definitions, the mechanism decomposes into four layers:

1. A minimal data structure. The tool's input is the entire todo array; each item has exactly three fields:

json
{
  "todos": [
    {
      "id": "1",
      "content": "Run the build",
      "status": "in_progress"
    },
    {
      "id": "2",
      "content": "Fix any type errors",
      "status": "pending"
    }
  ]
}

The state machine has three states: pending, in_progress, completed. No priorities, no dependencies, no due dates — every fancy field would become maintenance burden for the model.

2. Full-overwrite semantics. Every TodoWrite call submits the complete new list (the tool returns an oldTodos/newTodos diff for display), not an incremental patch. This avoids state-sync bugs like failed diff application, at the cost of a few extra tokens per call.

3. The system prompt enforces discipline. The model is instructed to use this tool "VERY frequently" and to follow a set of hard rules:

  • If a task has more than 3 steps, is non-trivial, or the user gave multiple requests, a list must be created;
  • At any moment there may be only one in_progress item;
  • The moment an item is done, mark it completed immediately — no batching;
  • When blocked, don't mark done — instead add a new task describing what needs resolving;
  • Tests failing, implementation incomplete, unresolved errors — none of these may be marked completed;
  • Tasks no longer relevant should be removed entirely to keep the list clean.

4. A single source of truth. The list lives in the conversation context, visible to the model at every generation turn — that's the anti-drift mechanism: no matter how much tool output piles in between, "where am I, what's left" is always at the most recently visible position.

Why "only one in_progress" is an important engineering decision

On the surface this constraint keeps the UI clean; the deeper effect is forcing serial focus on the model. LLMs on long contexts tend to multitask: mid-edit on file A, suddenly remembering problem B, hopping over to B, and coming back to find A's context has gone fuzzy. Forcing a single focus constrains the agent's behavior into one traceable trajectory — and when you debug afterwards (via observability logs), you can answer precisely "where did it go off the rails."

Why it works: cognitive offloading, not added intelligence ​

Note that TodoWrite has zero runtime logic — the harness doesn't check the list, doesn't enforce it, doesn't alert when the model forgets to update it (which is exactly what open-source reimplementations have run into: models frequently forget to update status; see opencode issue #28961). Its entire mechanism of action is cognitive offloading:

  1. The act of generating the list forces the model to think through the task's structure before touching anything;
  2. The list persists in the context, anchoring every decision round;
  3. The status-update action turns "review progress" into a fixed rhythm of the loop — every time it marks something completed, the model implicitly has to answer "did the last step actually finish? What's next?"

This distinction matters: the todo list isn't a progress bar for the user (though it serves that too) — it's external memory the model writes for itself. Which is also why it fundamentally differs from memory systems: it doesn't persist beyond the session and doesn't aim for retrievability — only for "always visible within the current session."

The granularity question ​

Explicit planning introduces a new problem: how fine should tasks be decomposed?

Too coarse ("implement the whole feature" as one task): the list degenerates into decoration, losing its tracking and course-correcting value — mid-task drift happens anyway.

Too fine (one item per function, per edit): the list itself eats context and tool calls; worse, a fine-grained plan is written when information is scarcest, so much of it is guaranteed to be invalidated during execution — the model ends up "maintaining the list" instead of "doing the task."

The engineering rules of thumb:

  • Use "verifiable checkpoints" as the unit. A good task has an objective completion criterion — "get the build passing and fix the type errors" is good; "change the code" is bad. This echoes TodoWrite's discipline directly: "tests red → no marking completed."
  • Coarse first, then refine on a rolling basis (rolling decomposition). The first draft lists only phase-level tasks (3–7 items); when entering a phase, expand it into substeps on the spot. Claude Code's own example works this way: first create four items ("survey the codebase → design → implement → export the feature"), and only after the survey completes does "fix the 10 type errors" get expanded into 10 items.
  • Never nest deeper than two levels. Three or more levels of task nesting is almost always over-engineering — if you truly need that depth, the decomposition belongs in subagents, not the todo list.

When to replan ​

Plans go stale; the question is what mechanism the harness uses to notice and respond. Common trigger signals:

TriggerExampleAppropriate response
The same action fails repeatedlyThe same test stays red after 3 fix attemptsHalt execution; re-examine the current task's assumptions
Observation contradicts the planThe plan assumed an ORM; the code is raw SQLRevise the current and downstream tasks
New subtasks discovered mid-executionWhile changing the API, discover a data migration is also neededInsert a new task into the list (micro-replanning)
External change in environment/requirementsThe user changes requirements midwayRetire the affected tasks; rebuild the tail of the list

Two negative lessons matter just as much:

  • Don't replan every step. BabyAGI-style "regenerate the task queue after every action" makes the plan oscillate between steps; the agent spends all its time adjusting the plan instead of working, and burns a pile of tokens. Replanning should be event-driven (failure, contradiction, new information), not time-driven.
  • Micro-revisions beat wholesale replanning. Cross one item off, insert two, reword one — in interleaved mode, the vast majority of "replanning" is exactly this kind of cheap operation. Only when the goal itself changes or the plan's overall assumptions collapse is it worth rewriting from scratch.

A subtle failure mode

The most dangerous replanning scenario is the model quietly abandoning the original plan without updating the list: the list still says "fix the type errors" while the model is off doing something else. The user sees a plan disconnected from actual behavior — worse than no plan at all, because it provides false auditability. This is precisely why disciplines like "mark completed immediately" and "add a task when blocked" get hammered into the system prompt.

The reasoning-model era: is explicit planning still necessary? ​

On September 12, 2024, OpenAI released o1-preview (with the official o1 following on December 5, 2024); on January 20, 2025, DeepSeek released and open-sourced R1. These reasoning models produce long chains of thought before answering, handling decomposition, self-checking, and backtracking internally. A natural objection follows: if the model can "think" on its own, why does the harness still need explicit planning?

It helps to separate two kinds of "planning":

What reasoning models internalize is reasoning — how to work through this piece of logic, how to fix this bug. It happens inside a single generation's chain of thought, dissipates when the output ends, and knows nothing about environment state.

What the harness externalizes is state — which items in the overall task list are done, which aren't, where things are stuck, what new situations surfaced midway. It must persist across multiple generations and multiple tool calls, which no single chain of thought, however long, can do.

From that distinction, a well-grounded judgment: reasoning models make ReAct-style "think step by step" redundant, while making TodoWrite-style "externalized state" even more necessary. Four reasons:

  1. Chains of thought don't persist. At call 30 the model can't see the chain of thought from call 2 (and even if it did, it would be truncated), but a todo list, as ordinary conversation content, stays visible throughout. Cross-turn state continuity can only come from externalization.
  2. Chains of thought can't be audited or interrupted. A user can't step in mid-"thought" to correct course; a todo list and plan mode's approval gate give human–machine collaboration a physical handle. Anthropic makes the same point in Building effective agents: agent systems trade latency and cost for task performance, so they need checkable, interruptible structure at critical junctures all the more.
  3. Chains of thought can't manage the environment. Plans fail primarily because environment feedback doesn't match expectations, not because the reasoning wasn't deep enough. However strong the model's reasoning, it can't know the world changed without re-observing it.
  4. The empirical evidence. After plugging in stronger reasoning models, systems like Claude Code and OpenHands did not remove TodoWrite / planning tools — they kept tightening the discipline around them. That's the most convincing industry vote.

A reasonable prediction: as models get stronger, planning will converge further toward "lightweight, structured, approvable" — fewer pre-set workflows, more model autonomy, but the externalized plan artifact will persist long-term as both the human–machine interface and the state anchor.

Practical guidance: when to force planning, when to let it freelance ​

If you're designing a harness or writing the usage contract for an agent, here's a decision checklist you can use as-is.

Force explicit planning (build a todo list / enter plan mode first) when:

  • The task contains 3+ identifiable steps (Claude Code's official threshold — a reasonable empirical value in practice);
  • The user gave multiple independent requests at once;
  • The task spans multiple files/modules/systems, where omissions creep in easily;
  • There are irreversible operations (database migrations, deletions, deployments) — the plan must be human-approved first;
  • It's a long session, expected to exceed a dozen or so tool calls;
  • The scenario requires handoffs between people or agents — the list is the handoff document.

Let it freelance (purely reactive) when:

  • Single-step or trivial tasks ("run the tests," "what does this function do");
  • Exploratory research — the goal is "figure out what's going on," and there isn't even a basis for decomposition yet; writing a plan here is pure ritual;
  • Highly uncertain tasks where every step depends on the previous observation (interactive debugging, penetration testing) — a pre-set plan has a half-life measured in minutes;
  • Pure conversation / information lookup.

A meta-rule

The criterion isn't "how big is the task" but "how long does the plan stay valid." Long validity (certain steps, stable environment) → explicit planning is worth it. Short validity (exploration, debugging, adversarial environments) → reactive + rolling refinement. When in doubt, default to: build a coarse list (3–7 items) and refine as you go.

Further reading ​

References ​