Skip to content

Planning and Task Decomposition

At a glance A systematic tour of the four classic agent planning patterns (ReAct, Plan-and-Execute, Tree of Thoughts, Reflexion), the engineering value of todo-list mechanisms, the debate over explicit planning in the reasoning-model era, and countermeasures for failure modes like over-planning and plan rigidity.

Planning and Task Decomposition ​

An agent that only "takes one step and looks one step ahead" looks fine in single-turn Q&A; stretch the task to dozens of steps across multiple tools and it loses the goal, circles aimlessly, and declares completion halfway through. The planning module exists to fix exactly that: turning "what to do next" from the model's improvisation into an explicit process that can be checked, revised, and measured.

From an engineering perspective, this piece maps the full spectrum of planning capability: what problem each of the four classic patterns solves; why the todo list became standard equipment on mainstream coding agents in 2025; whether explicit planning still has a reason to exist now that reasoning models have risen; and when a planning module is warranted versus when it's pure over-engineering.

1. Where Planning Sits in an Agent ​

Recall the classic loop from the agent system anatomy: perceive → decide → act → observe. Without planning capability, the "decide" step is purely reactive — the model looks at the current context and improvises the next action. That's the pre-ReAct shape of agents, and the shape of many toy demos.

Reactive decision-making has three structural weaknesses:

  1. Goals dilute. As observations keep flowing into the context window, the original user goal carries less and less attentional weight, and the agent gets swept along by intermediate steps.
  2. Errors compound. Every step improvises on the previous step's result; once step 3 goes sideways, step 4 doubles down in the wrong direction, and nobody checks back.
  3. Resources can't be budgeted. The user can't predict how many tokens the task will burn, how many tools it will call, or how long it will run — so there's no way to intervene mid-flight.

The essence of planning is adding a layer of slow thinking above the reactive loop: think through (or write down) "what to do" first, then let the fast loop execute "how to do it." In the language of the Agent Loop, planning is a metacognitive process that runs outside the loop (or at specific nodes within it).

Agent without planning (reactive):
  goal ──> [observe → decide → act] ──> [observe → decide → act] ──> ...
               ↑ every step improvised; nobody notices when it goes wrong

Agent with planning:
  goal ──> [Planner: generate a plan] ──> [Execute step 1 → 2 → 3 ...]
                                               ↑
                                    [Replanner: drifting? revise the plan]

2. The Spectrum of Classic Patterns ​

Planning research isn't a single line; it's four complementary paradigms, each aimed at a different failure point. Understanding their differences beats memorizing the names.

2.1 ReAct: think while doing (2022) ​

ReAct (Reasoning + Acting, Yao et al., arXiv:2210.03629, ICLR 2023) is the foundational pattern of all modern agents. Its insight is plain: at each step, have the model first emit a Thought (reasoning), then an Action (tool call), take the Observation, and continue. The reasoning trace and the action trace interleave and anchor each other — Thoughts give Actions grounds; Observations keep Thoughts honest.

python
# Minimal ReAct loop skeleton (pseudocode; shows the prompt structure)
messages = [system_prompt, user_task]
while True:
    reply = llm(messages)  # the model outputs Thought + Action
    if reply.has_tool_call:
        observation = execute_tool(reply.tool_call)
        messages.append(reply)            # Thought: ... Action: search("...")
        messages.append(observation)      # Observation: search results...
    else:
        break  # the model emitted the final answer; the loop ends

ReAct's planning is local and immediate: each Thought covers only the step at hand. On short tasks that's a strength (flexible, no wasted tokens); on long tasks it's the source of all three weaknesses in Section 1. Today, nearly every coding agent's (including Claude Code's) main loop is still ReAct-shaped — the difference is only in what's wrapped around it.

2.2 Plan-and-Execute: plan first, execute second (2023) ​

Plan-and-Solve Prompting (Wang et al., arXiv:2305.04091, ACL 2023) first proved on pure-reasoning tasks that having the model draft a plan and then solve subtasks one by one significantly reduces the "skipped steps" errors of zero-shot CoT. That same year LangChain engineered the idea into the Plan-and-Execute pattern, and early autonomous agents like BabyAGI took a similar route.

The architecture draws a clean role separation:

User task
   │
   ▼
┌──────────┐   full step list   ┌───────────┐
│ Planner  │ ────────────────▶  │ Executor  │ ──▶ calls tools step by step
│ (large   │                    │ (a smaller│
│  model)  │                    │  model OK)│
└──────────┘                    └─────┬─────┘
      ▲                               │ execution results
      │        ┌────────────┐         │
      └────────│ Replanner  │◀────────┘
               │ revise the plan on deviation│
               └────────────┘

Compared with ReAct, it buys three things: the global goal stays written in the plan, immune to being diluted by observations; the Planner and Executor can run on different model tiers, saving money; and the plan itself is a structured artifact that can be persisted and shown to the user for confirmation. The price is rigidity — the plan is fixed before execution begins, and once reality diverges from expectations, you either trigger replan constantly (slow and costly) or grind through an obsolete plan.

The extension along this line is LLMCompiler (Kim et al., arXiv:2312.04511): the Planner's output isn't a linear step list but a dependency-aware DAG, with independent tool calls executing in parallel. The paper reports up to 3.7× latency speedup and 6.7× cost savings versus ReAct. The idea is elegant but narrow — it only makes sense when tool dependencies can be determined in advance.

2.3 Tree of Thoughts: search-style planning (2023) ​

Both previous patterns "walk one road to the end." Tree of Thoughts (Yao et al., arXiv:2305.10601, NeurIPS 2023) changes the frame: since a single reasoning chain can dead-end, turn reasoning into search — generate multiple candidate Thoughts per step, have the model score each candidate's value (sure / maybe / impossible), expand the search tree with BFS/DFS, and backtrack when a branch dies.

Problem: make 24 from 4 numbers (Game of 24)
                    ┌─ Line A: make 6 and 4 first ── scored: maybe ── expand ─┐
  Root ── generate 3 ┼─ Line B: make 8 and 3 first ── scored: sure   ── expand ─┤── ... ── solution found
  candidates        └─ Line C: make 12 and 2 first ─ scored: impossible ✗ pruned

ToT is stunning on tasks like Game of 24 where the solution space needs exploring, but stay clear-eyed: its evaluator is still the LLM itself, so evaluation quality caps search quality; and every node is a model call, so cost explodes with branching. In agent engineering, pure ToT is rarely run in production directly — its ideas were mostly absorbed into the training of reasoning models (see Section 4).

2.4 Reflexion: reflect and correct (2023) ​

Reflexion (Shinn et al., arXiv:2303.11366, NeurIPS 2023) targets a different failure point: the mistakes an agent made last round get repeated next round, because the trajectory vanishes once it passes. Reflexion adds a layer of verbal reflection after task failure — the model writes a natural-language postmortem of the failed trajectory ("what went wrong this time, what to change next time"), stores it in episodic memory, and brings it into the context on the next attempt.

Attempt 1: execution trajectory ──▶ failure signal (tests failed / environment feedback)
                                      │
                                      ▼
                            Reflector: generate the reflection text
                            "I assumed the API returns a dict, but it's a list;
                             next time print the type before parsing"
                                      │
                                      ▼
                            Store in memory ──▶ Attempt 2's prompt carries the reflection ──▶ passes

The paper reports it pushed GPT-4's pass@1 on HumanEval to 91%. Note the contrast with ToT: ToT searches within a single task, Reflexion learns across attempts; ToT prevents wrong turns, Reflexion cures repeated stumbles. Engineering-wise Reflexion is dirt cheap — no training, just one extra LLM call plus a piece of text memory — making it one of the highest-ROI planning enhancements. Its boundary with the memory system also sits here: the reflection text is essentially forward-looking experiential memory.

How to choose

The four patterns are not substitutes; they're orthogonal tools. The common production recipe: ReAct as the main loop + a todo list for lightweight planning (Section 3) + one reflection pass triggered on failure. Pure Plan-and-Execute suits long flows with predictable steps (data pipelines, report generation); save ToT for offline tasks that genuinely require exploration.

3. The Todo-List Mechanism: an Underrated Planning Primitive ​

Since 2025, frontline coding agents like Claude Code and Codex have all shipped todo_write-style tools. That's not coincidence — it's a very well-priced engineering decision. Someone reverse-engineered Claude Code's actual API requests (see References), and several details are worth copying:

  • The todo is usually the first tool call. Get a complex task → build the list → then start work.
  • Every update rewrites the whole list. No incremental-update logic, full overwrite — crude, but it eliminates inconsistent state.
  • The tool result carries instructions. TodoWrite's tool result always appends a reminder like "continue to the next item on the list" — reinforcing the behavior on every call, with far higher compliance than writing it once in the system prompt.
  • The system injects reminders dynamically based on list state. For example, when the list is empty it nudges "consider creating todos"; after a change, it pastes the latest list into the conversation.

Why does one todo table work this well? Because it solves three problems at once:

  1. It fights goal dilution. Task state moves from "buried in hundreds of turns of history" to "a permanently visible table in the context"; every glance at the list re-anchors the goal. This is really a context engineering trick: put the most important state in the most prominent position.
  2. Progress becomes observable. The list doubles as a UX component — the user can see which step the agent is on, which is cheap observability.
  3. Intervention gets a handle. The user can directly say "skip item 3" or "do the last item first" — the plan becomes a shared artifact of human-agent collaboration, a natural fit with Human-in-the-Loop design.

The todo list can be seen as Plan-and-Execute's "economy class": the plan is no longer a contract fixed before execution, but a living document rewritten continuously during execution. It gives up strictness (no forced replan triggers) in exchange for zero-friction compatibility with the ReAct main loop. Practice shows that for most coding and office tasks, this compromise beats a strict two-tier architecture.

Don't over-trust the list

The todo mechanism guards against "forgetting," not against "doing wrong." Every item on the list can be executed incorrectly, and the agent can mark unfinished items as completed. The list is scaffolding for memory, not a guarantee of correctness — acceptance still depends on external signals: tests, lint, human review.

4. The Reasoning-Model Era: Has the Long Chain of Thought Replaced Explicit Planning? ​

September 2024 brought OpenAI's o1; January 2025 brought the open-sourcing of DeepSeek-R1, and the "reasoning model" paradigm was established: models trained via reinforcement learning to produce internal long chains of thought, reasoning for tens of seconds to minutes before answering, spontaneously decomposing, verifying, and backtracking along the way — sounding like CoT, Self-Consistency, ToT, and Reflexion all internalized into the weights.

So a debate ran through 2025–2026: can explicit planning modules retire to the museum?

The honest answer: they replaced "reasoning-type planning," not "engineering-type planning."

What was replaced: puzzle-style planning. Math proofs, logic riddles, single-file algorithm problems — for these, planning happens entirely inside the model's head, needing no tools and no intermediate artifacts. Reasoning models genuinely make hand-written ToT searches redundant here; the carefully designed prompt tricks of papers got absorbed directly by training.

What wasn't replaced, for very concrete reasons:

  1. The plan's audience isn't just the model. A long chain of thought is a private monologue, discarded after running; an explicit plan is a public artifact that can be shown to the user for confirmation, written into a ticket system, and read by another agent. In multi-agent collaboration, the plan is the interface documentation (see multi-agent architecture).
  2. Long tasks exceed a single inference's span. A few minutes of thinking cannot govern a task that runs for hours. Cross-session task state must be externalized — the very reason todo lists and memory systems exist.
  3. Verifiability and accountability. When an agent errs, an explicit plan gives evaluation and debugging an anchor: was the plan wrong or the execution? Against a black-box chain of thought, that question is nearly unanswerable.
  4. Controllable cost. Reasoning models bill for thinking tokens, unpredictably; explicit planning can concentrate "thinking it through" into one high-quality call, with execution on a cheaper model — Plan-and-Execute's cost-structure advantage actually sharpens in the reasoning-model era.

The 2025–2026 consensus in engineering and research (see Zylos's survey in the References) is that the #1 problem for agents failing in deployment isn't shallow reasoning but plan rigidity — inability to backtrack and replan when reality forks from expectation. That's precisely what built-in reasoning-model capability doesn't cover and only architecture can solve. Hence today's best practice is layering: reasoning models handle "think each step through," explicit mechanisms handle "never lose the thread across the whole run."

5. Hierarchical Decomposition and Subgoal Management ​

When a task outgrows a flat todo table (dozens to hundreds of subtasks), you need hierarchical decomposition: high-level goal → milestones → subtasks → atomic actions. This is the classic HTN (Hierarchical Task Network) planning idea from traditional AI, with three engineering points in the LLM-agent era:

Decompose to the "verifiable" boundary. A subtask should stop splitting once it can be judged complete by one clear signal — tests pass, the API returns 200, the file is generated and format-validated. Splitting further ("open the file," "read line 10") wastes planning budget; that territory belongs to execution-layer improvisation.

Declare dependencies between subgoals explicitly instead of assuming order implicitly. A flat list defaults to sequential execution, but "design the database schema" and "write the frontend page" can actually run in parallel, while "integration testing" must wait for both. Write dependencies down, and you know the blast radius when something fails and what can be retried. LLMCompiler's DAG idea applies equally well to human-designed workflows.

High-level plans are stable; low-level plans are volatile. In a good hierarchy, the closer a goal is to the root, the less it should change (if it changes, the requirement itself changed); the closer to the leaves, the more freely it can be overturned. Constraining "replanning" to the lowest possible level is the key to controlling replan cost — don't regenerate the whole project plan because one API call failed.

A typical three-level structure:

Goal: add an "export report" feature to the product
├─ M1: backend export endpoint        ← milestone; clear acceptance: the endpoint returns a correct xlsx
│   ├─ Design the data model for the export job
│   ├─ Implement the generation logic (depends on: data model done)
│   └─ Endpoint + unit tests (depends on: generation logic done)
├─ M2: frontend export button + progress polling  ← can be developed in parallel with M1
└─ M3: integration and acceptance     ← depends on M1+M2

In practice, it's worth doing the top-level decomposition once, seriously, with the strongest model (even having a human review it), then handing lower-level subtasks to executor agents to run freely in ReAct loops. This is also the idea behind Claude Code's Task tool dispatching sub-agents: the main agent owns decomposition and acceptance, sub-agents own execution; neither knows the other's details, and they interface through task descriptions and acceptance criteria.

6. Planning Failure Modes and Countermeasures ​

Planning modules introduce their own failure modes. In 2025–2026 production retrospectives, these four show up most:

Failure modeSymptomsRoot causeCountermeasures
Over-planningEven simple tasks burn thousands of tokens on a plan first; planning takes longer than executionApplying the planning template regardless of task complexitySet a complexity gate (Section 7); let the agent judge "this task needs no plan"
Plan rigidityThe environment changed (an API changed, a file is gone) but the original plan runs to the end anywayMissing replan triggers, or trigger thresholds too highHard rule: any step failing N times, or any observation contradicting the plan's assumptions, must stop and replan
Goal driftMid-task it starts optimizing something irrelevant; the deliverable no longer matches the original requestThe original goal diluted in a long contextMake goals explicit (first todo item / a per-turn reminder restating the goal); re-read the original request at milestone acceptance
Subtask explosionPlans recursively spawn sub-plans, levels deepen, tokens burn with no outputDecomposition lacks a termination condition; "planning" is mistaken for progressCap decomposition depth (2–3 levels max); restrict "only subtasks over X steps may split further"; monitor planning-token share

Two cross-cutting engineering disciplines:

  • Budget the planning itself. If planning consumes more than 20–30% of total budget in tokens/time without execution having started, something is wrong — either the task should be split off to a human, or the planning strategy should be downgraded.
  • Failure signals must come from the environment, not the model's self-assessment. "I feel this step is done" doesn't count as done. Tests, schema validation, HTTP status codes, file diffs — only steps with access to external signals have earned the right to be marked completed. Fail this, and the todo list degenerates into a self-congratulation list.

For more cautionary tales, see the field guide to pitfalls.

7. Practical Advice: Which Tasks Deserve a Planning Module ​

Finally, an actionable decision framework. Split along two dimensions: step count and step predictability:

                Steps highly predictable      Steps poorly predictable
             ┌─────────────────────┬─────────────────────┐
  Few steps  │ No planning module   │ Bare ReAct loop       │
  (<5 steps) │ execute directly /   │ improvised decisions  │
             │ fixed workflow       │ at each step are fine │
             ├─────────────────────┼─────────────────────┤
  Many steps │ Plan-and-Execute     │ ReAct + todo list      │
  (>10 steps)│ strict plan + DAG    │ + Reflexion triggered  │
             │ parallelism          │ on failure             │
             ├─────────────────────┼─────────────────────┤
  Exploratory│ Tree of Thoughts /   │ Hierarchical decomposition│
  tasks      │ sample multiple      │ + sub-agents; main agent│
             │ approaches, then pick│ owns decomposition &    │
             │                      │ acceptance              │
             └─────────────────────┴─────────────────────┘

Three compressed lessons:

  1. For tasks under 5 steps that one person could do in one pass, any planning module is over-engineering. Run the ReAct loop directly and save the budget for execution.
  2. For cross-tool tasks over ten minutes whose steps can be roughly listed, add a todo list first — the highest-ROI single move, and often the only one needed. Copy the four implementation details from Claude Code in Section 3.
  3. Adopt hierarchical decomposition + sub-agents only when the task can be split into parallel subtasks with clear acceptance criteria AND failure is expensive. The debugging and maintenance cost of that architecture is real; most applications don't need it.

If you want to build a planning-capable agent from scratch, follow the route in the build-your-own agent tutorial: implement the ReAct main loop first, then add the todo tool, then experiment with failure reflection — each step lets you watch success rates and token costs move. For deeper readings of the related papers, see core papers in depth.

References ​