Appearance
The Agent Loop
If the entire agent field could be remembered as one picture, it would be this one: the model outputs → an action executes → the result is observed → fed back to the model, over and over, until the task is done. This cycle is the Agent Loop, and it's the watershed between "chatbot" and "agent" — the former is a function call; the latter is a stateful process.
This is the most important page on the site. It answers four questions: why a single LLM call isn't enough; what actually happens inside the loop (with pseudocode you can copy and adapt); when the loop should stop (much harder than it sounds); and what differences in loop design are worth stealing from real products — Claude Code, SWE-agent, OpenHands. It closes with a countermeasures table for runaway patterns, something you'll run into daily once you're debugging agents.
1. Why a Loop: a Single Call Can't Do Multi-Step Tasks
Start with a plain question: why can't one call just finish the job?
Suppose the task is "fix the 500 error on this repo's login endpoint." A strong enough model could, in theory, emit the complete patch in one shot. In practice that road doesn't work — not because the model isn't smart, but because the information isn't present:
- The model doesn't know what the login endpoint's code looks like, so it must read the file;
- After reading, it discovers the error comes from a downstream database query, so it must read another file;
- It changes the code but doesn't know whether the change is right, so it must run the tests;
- The tests fail, and the failure output determines whether the next move is fixing the code or fixing the test.
What information each step needs depends on what the previous step saw. The goal is known up front, but the path is not — that is the very reason loops exist: defer the "what do I do next" decision until after the latest observation arrives, and let the model converge on the answer through real environmental feedback instead of betting on a complete plan with incomplete information.
Anthropic draws this distinction cleanly in Building Effective Agents: a workflow orchestrates LLM calls through predefined code paths (the path is hard-coded); an agent has the LLM dynamically direct its own process and tool usage (the path is decided live by the model). The Agent Loop is the runtime vehicle of the latter.
A quick test
If your task's path can be drawn completely as a flowchart — what comes first, what comes after, how every branch is handled, all writable in advance — what you want is a workflow, not an Agent Loop. The loop's value lies precisely in paths that can't be drawn. A large share of failed "agent projects" crammed workflow-shaped work into a loop, paying unpredictable cost and latency without buying any flexibility.
2. Anatomy of the Classic Loop: Perceive → Think → Act → Observe
Break the loop open and each iteration walks four fixed beats:
┌─────────────────────────────────────────┐
│ │
▼ │
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ Perceive │──▶│ Think │──▶│ Act │──│ Observe │──┘
└──────────┘ └──────────┘ └──────────┘ └──────────┘
Assemble the LLM reasons Parse & run Results fed back
context & decides the tools into the context- Perceive: no actual sensors involved. For an LLM agent, "perceiving" means assembling the context sent to the model this turn: system prompt, tool catalog, task description, message history, the previous step's observations. This beat is pure engineering, and it's the home turf of context engineering.
- Think: call the LLM. Based on the current context, the model emits a decision — possibly reasoning plus a tool call (ReAct style), possibly structured
tool_useblocks (native function calling style). This is the loop's only "intelligent" beat; the other three are deterministic code. - Act: parse the model's chosen action out and execute it in the real environment — run shell commands, read and write files, call APIs. This is where you handle all the dirty cases: illegal model output, nonexistent tools, execution timeouts.
- Observe: format the execution results (stdout, errors, file contents, return values) into a message and append it to the message history. The next iteration's "perceive" reassembles it into the context. Observations are the loop's fuel — at bottom the loop is a context-accumulation process driven by observations.
Of the four beats, thinking and acting together are often called "the model's move," while perceiving and observing are the harness/scaffold's job. A fact the industry internalized after 2024: the same model wearing different harnesses produces wildly different scores — SWE-agent took an untuned GPT-4 Turbo to a 12.5% SWE-bench solve rate purely through interface design, more than triple the previous best of 3.8%. The harness isn't wrapping paper; it's combat power.
3. Complete Agent Loop Pseudocode
The pseudocode below covers every component a production-grade loop needs: context assembly, model invocation, tool parsing and execution, observation feedback, stop conditions, a step cap, and error fallback. It's written Python-style, with syntax aligned to the Anthropic Messages API's tool-use flow (tool_use / tool_result blocks, stop_reason), but it's bound to no framework — indeed Anthropic itself recommends this: many patterns take just a few dozen lines of raw API code, and understanding the layer beneath the framework matters more than the framework.
python
MAX_STEPS = 40 # hard step cap: the fuse against runaway loops and runaway cost
MAX_CONSEC_ERRORS = 3 # circuit-breaker threshold for consecutive errors
def agent_loop(task: str, tools: list[Tool]) -> str:
# ── 1. Initialize the message history (the loop's only state) ─────────────
messages = [
{"role": "system", "content": build_system_prompt(tools)},
{"role": "user", "content": task},
]
consec_errors = 0
# ── 2. Main loop ─────────────────────────────────────────
for step in range(MAX_STEPS):
# 2a. Perceive: assemble the context (compress/trim here if needed — see context engineering)
context = assemble_context(messages)
# 2b. Think: call the model, passing the tool catalog
response = llm.create(
model="claude-sonnet-4-6",
messages=context,
tools=[t.schema for t in tools], # name + description + input_schema
)
# Append the model's raw output to the history (keep the reasoning on record too)
messages.append({"role": "assistant", "content": response.content})
# 2c. Stop check: no tool calls = the model believes it's done
if response.stop_reason == "end_turn":
return extract_text(response) # extract the final answer, exit the loop
# 2d. Act: parse and execute each tool call
tool_results = []
for block in response.tool_calls():
try:
tool = find_tool(tools, block.name) # raises if the tool doesn't exist
result = tool.execute(block.input) # real execution (with a timeout)
consec_errors = 0 # one success resets the error counter
tool_results.append(tool_result_block(block.id, result))
except Exception as e:
# Key design: don't throw errors out of the loop — feed them back as observations.
# Letting the model see the failure and correct itself is the core
# fault-tolerance mechanism of agents.
consec_errors += 1
tool_results.append(tool_result_block(
block.id, f"Error: {e}", is_error=True))
# 2e. Observe: execution results enter the history as a user message
messages.append({"role": "user", "content": tool_results})
# 2f. Error circuit breaker: consecutive failures mean the model is stuck in a
# situation it cannot resolve
if consec_errors >= MAX_CONSEC_ERRORS:
return f"{consec_errors} consecutive tool failures — escalating to a human."
# ── 3. Steps exhausted: the fuse has blown ────────────────────
return f"Reached the max of {MAX_STEPS} steps; task incomplete — needs human review."Several design decisions in this pseudocode deserve a closer look:
- State is the message history, nothing else. The agent's entire memory, progress, and intermediate conclusions live in the
messagesarray. That makes the loop naturally serializable, resumable, and auditable — persist the array and you have a complete trace. - Errors are observations, not exceptions. Tool failures (command not found, invalid arguments, execution timeout) get wrapped as
is_error=Truetool results and fed back to the model. On seeing the error, the model can usually self-correct — a mechanism validated back in the ReAct paper and the main source of modern agents' robustness. But consecutive failures must trip the breaker, or you have an expensive infinite loop. - Stopping relies on
stop_reason, not string matching. No tool call in the model's output (manifested asstop_reason == "end_turn"in the Anthropic API) means it has decided to wrap up. This is the dominant stop signal of the function-calling era — next section covers it in detail. - A step cap is mandatory, not optional. A loop without
MAX_STEPSis a car without brakes. There's a real case from August 2026 in the ollama repo: a model fell into a self-sustaining loop against a Claude Code-compatible endpoint, emitting 193 nearly identical tool calls in a row and burning roughly 31 million input tokens before a human interrupted it.
The distance between pseudocode and production
This code works, but a production loop needs at least: compression/truncation when the context overflows (see Context Engineering), structured per-step logging and replay (see Observability), cost metering and budget breakers (see Cost Engineering), and human-approval gates for dangerous operations (see Human-in-the-Loop). The loop itself is simple; making the loop reliable is the hard part.
4. The Art of Stop Conditions
"When to stop" is the most underestimated problem in Agent Loop design. Stop too early and the task is unfinished; stop too late and it burns money or causes damage. Practice has converged on four stop mechanisms, usually combined in production.
4.1 The Model Declares Completion (end_turn)
The default mechanism of the function-calling era: when the model stops emitting tool calls and outputs only natural language, the loop ends. Pros: zero cost, zero intrusion. Cons: the model's "I'm done" and "it's actually done" are frequently not the same thing — the model may have forgotten the original request because the context got long, capitulated early after repeated setbacks, or simply gotten lazy. It answers "does the model want to stop," not "is the task complete."
4.2 A Structured finish Tool
Give the model an explicit finish (or submit, done) tool, with the convention that only calling it ends the task. SWE-agent is the textbook case: the model must call the submit command to hand in its patch before the loop terminates. The benefits: completion becomes a concrete, parameter-carrying action (submit the final answer, the patch path), and you can put a verification gate inside finish's execution logic — automatically run the tests before accepting; if they fail, push the failure back as an observation and refuse to close. That upgrades "the model says so" into "the model says so + the environment confirms" — the most effective known cure for premature exits.
4.3 Budget Exhaustion (steps / tokens / time / money)
The hard-constraint fuse: MAX_STEPS is only the most common variant. Production systems usually stack several: max iterations, max token spend, max wall-clock time, max dollar amount. SWE-agent treats budget control as a first-class citizen and aborts when the cost ceiling is hit. Think hard about the exit behavior when the budget dies: dumping whatever exists is the worst option; the good pattern is to assemble a handoff report — current progress, confirmed conclusions, where it got stuck — and exit with that.
4.4 Stalemate Detection (loop detection)
The model didn't say stop, the budget isn't exhausted, yet the loop is already dead — the last N steps emit the same or near-same actions, observations stop changing, context growth stalls. These "zombie loops" evade all three mechanisms above and need explicit detection: compare the similarity of the last K actions (simplest form: hash dedup), and on hitting the threshold interrupt or degrade (e.g. force the model to summarize the situation before deciding). The 193-identical-calls incident mentioned above would have been caught by a three-step repeat detector.
| Stop mechanism | The question it answers | Strengths | Risks |
|---|---|---|---|
| Model declares completion | Does the model want to stop? | Zero cost, natural | Premature exits, laziness |
| finish tool + verification | Is it actually done? | Completion is verifiable and carries artifacts | Verification logic must be right, or it misjudges |
| Budget exhaustion | Can we still afford it? | Reliable backstop | Interrupting loses progress (needs a handoff report) |
| Stalemate detection | Is the loop still alive? | Zombie-loop killer | Similarity threshold can punish legitimate retries |
5. The ReAct Pattern, Closely Read
The Agent Loop is the skeleton; ReAct is the decision-making pattern that gives the skeleton a soul. It comes from the paper "ReAct: Synergizing Reasoning and Acting in Language Models" (Shunyu Yao et al., Princeton University and Google Brain, posted to arXiv in October 2022 as arXiv:2210.03629, later published at ICLR 2023 as an oral). Built on PaLM-540B, the paper systematically validated the approach on two knowledge-QA benchmarks (HotpotQA, FEVER) and two interactive-decision benchmarks (ALFWorld, WebShop).
5.1 The Core Mechanism: Three Fields, Alternating
ReAct's insight is so plain it sounds obvious in hindsight: have the model emit a reasoning trace (Thought) at each step, then an Action, with the environment returning an Observation — the three interleaved:
Thought 1: I need to find the login route definition first; it should be under app/routes.
Action 1: search_dir["login", "app/routes"]
Observation 1: Found 3 files: auth.py, session.py, oauth.py ...
Thought 2: auth.py looks most relevant — read it, focusing on the database query parts.
Action 2: view["app/routes/auth.py"]
Observation 2: (file contents... line 47: db.query(User).filter_by(...))
Thought 3: The filter_by call on line 47 will throw when username is None...
Action 3: edit["app/routes/auth.py", ...]Why interleaved rather than separate? The paper's argument: Thoughts give Actions grounds (each action derives from reasoning rather than blind guessing), and Observations keep Thoughts honest (reasoning is continually corrected by real environmental feedback). Pure reasoning (Chain-of-Thought) tends to hallucinate a wrong premise when facts are missing and snowball from there; pure action (Actions without Thoughts) lacks task decomposition and exception handling. The two working together anchor each other.
5.2 What ReAct Looks Like Today
A common misconception: "ReAct is obsolete; everything is native function calling now." The opposite is true: function calling is ReAct's industrialized packaging. Today's tool_use mechanisms in Claude, GPT, and other models essentially fix the Thought (text/thinking blocks in the assistant message) + Action (tool_use block) + Observation (tool_result block) triple into the API structure, sparing you the prompt-format gymnastics. Open the trace of Claude Code or any modern coding agent and you'll find the original Thought/Action/Observation alternation intact. ReAct wasn't replaced; it was built in.
5.3 ReAct's Failure Modes
The original paper honestly documented its own failures. The most typical is reasoning-trace drift: in long chains, later Thoughts increasingly reason from previously generated Thoughts instead of from Observations, gradually detaching from the environment and self-reinforcing in the wrong direction. Many ALFWorld failures stem from this. The engineering countermeasures: periodically have the model do a "reality check" against the original goal and accumulated observations, or apply explicit planning-level course correction on long trajectories — which also explains why pure ReAct often needs a bolt-on planning module for long tasks.
6. Loop Design Compared Across Real Products
The theory of the Agent Loop is the same everywhere; the engineering differs enormously. Three of the most representative systems — their differences almost define the loop design space itself.
6.1 Claude Code: a Minimal Main Loop + Stop Means Ship
Claude Code's main loop is the closest of the three to Section 3's pseudocode: a while loop that calls the model repeatedly; as long as the response contains tool_use blocks it executes the tools, appends the tool_result, and continues — until the model returns text with no tool call (stop_reason == "end_turn"), at which point the loop ends and the text ships to the user. No explicit finish tool, no planner, no verification gate — the stop condition is entirely the model's own declaration.
The minimalism is deliberate. Anthropic says explicitly in Building Effective Agents that the most successful implementations tend to be simple, composable patterns rather than complex frameworks, and recommends developers build directly on the API and understand the layer beneath before reaching for frameworks. Claude Code moved the complexity elsewhere: tool design (small, crisp tools like Read/Grep/Edit), the system prompt, permission gates (dangerous commands need user approval), and CLAUDE.md project memory. The loop itself is paper-thin; what's thick is everything around the loop.
6.2 SWE-agent: Everything in Service of the ACI
SWE-agent (Princeton NLP, NeurIPS 2024, arXiv:2405.15793) has an equally plain loop — the model emits a command, the environment executes, results feed back — but it proved that the interface layer executing inside the loop matters a hundred times more than the loop itself. Its ACI (Agent-Computer Interface) is a set of commands and feedback formats designed specifically for LLMs, with core features including:
- Search result throttling: commands like
find_fileandsearch_dirreturn at most ~50 results per call, keeping observations from drowning the context; - A windowed file viewer: a stateful file viewer shows ~100 lines at a time with
scroll_up/scroll_downnavigation, rather than dumping whole files; - Inline-feedback edit commands: edit commands carry syntax checks — a broken edit errors immediately instead of surfacing at test time;
- An explicit
submit: the model must call the submit command to hand in its patch before the loop terminates (the finish-tool pattern from 4.2).
The evidence is hard: the same untuned GPT-4 Turbo, wearing this ACI, reached 12.5% on SWE-bench — more than triple the 3.8% of the best non-interactive method at the time — and the paper systematically ablated each ACI design's contribution. One-line conclusion: tool design beats model capability; within the loop's budget, spend first on the perceive and act beats.
6.3 OpenHands: the Event Stream as a First-Class Citizen
OpenHands (formerly OpenDevin, arXiv:2407.16741) abstracts the loop into three components: the Agent abstraction, the Event Stream, and the Agent Runtime. Its pivotal design upgrades "message history" into an event stream strictly ordered by time — every agent action and every environmental observation is a record in the stream; the agent assembles its prompt from the stream when deciding, and new events append after execution.
The engineering dividends of this abstraction:
- State is the event stream: the agent's entire state is serializable, replayable, and forkable, naturally supporting resume-from-checkpoint and human audit;
- Observer pattern: other components (UI, logging, evaluators) subscribe to the stream without intruding into the loop itself;
- Multi-agent delegation: sub-agent executions are modeled as events too (delegation action / observation), so the main loop's structure never changes.
OpenHands' default agent is CodeActAgent, following the CodeAct route: actions are not scattered JSON tool calls but composable Python code executed in a sandbox, with stdout/stderr fed back as observations.
6.4 The Three Compared
| Dimension | Claude Code | SWE-agent | OpenHands |
|---|---|---|---|
| Loop skeleton | while + tool_use, minimal | Single-command stepping, minimal | Event-stream driven, thickest abstraction |
| Stop condition | Model declares it (end_turn) | Explicit submit command | finish action / budget breaker |
| Action shape | Structured tool calls | Custom ACI commands | CodeAct: Python code |
| Observation design | Raw tool returns + truncation | Throttled/windowed, carefully curated | Modeled uniformly as events |
| Core bet | Strong model + thin harness | Interface design decides everything | Abstraction buys extensibility |
| Suited for | Interactive coding assistant | Academic benchmarks / batch jobs | General software-development platform |
None of the three loops is "correct" — the differences come entirely from the setting: interactive products need minimal latency, academic benchmarks need control and reproducibility, platforms need extensibility. When designing your own loop, first decide which row you stand in.
7. Common Runaway Patterns and Countermeasures
The Agent Loop's most notorious property: it can do work, and it can also burn through your entire budget at remarkable efficiency while accomplishing nothing. The three most frequent runaway patterns, plus field-tested countermeasures:
7.1 Infinite Loop
Symptoms: the model calls the same tool with the same arguments and gets the same result, forever. Typical triggers: the tool's error message never made it back into the context (the model doesn't know it failed), or the model is stuck in a "maybe this time it'll work" local optimum. The extreme case is the incident mentioned earlier — 193 identical calls, 31 million tokens.
Countermeasures: action-repetition detection (hash-dedup the last K actions; trip the breaker on a hit); feed tool errors back prominently with is_error; always keep both fuses — MAX_STEPS and a cost cap.
7.2 Thrashing
Symptoms: no strict repetition, but the loop oscillates between two states — changing file A breaks test 1, reverting breaks test 2, changing back breaks test 1 again… the trajectory keeps growing and progress stays at zero. Unlike an infinite loop, the action sequence doesn't repeat, so hash detection can't catch it.
Countermeasures: observation-level stagnation detection (if the last N observations are too similar, declare stagnation); on stagnation, force the model to stop and do a "situation summary + path reflection" before continuing, breaking the oscillation's momentum; set task-level milestones and escalate (change strategy, call a human) when a deadline passes unmet. These problems are fundamentally missing planning — see Planning & Task Decomposition.
7.3 Premature Exit
Symptoms: the task is half done and the model confidently outputs "complete" and calls end_turn. Triggers include: the original goal getting diluted in a long context, retreat after repeated setbacks, and the model's inherently lower bar for "done." This is the sneakiest of the three — the loop terminates normally; the result is just wrong.
Countermeasures: most effective is Section 4.2's verification gate — finish isn't the end; the environment's verification passing is. Second, at the prompt level, write the completion standard as a checkable list ("before submitting: all tests pass, changes cover requirements items 2 and 3"). For long tasks, anchor the goal with a todo-list mechanism. On the evaluation side, these failures only surface through end-to-end task success rates — a trace alone tends to look like "every step was reasonable." See Agent Evaluation.
| Runaway pattern | Observable signal | Root cause | First-line countermeasure |
|---|---|---|---|
| Infinite loop | Actions repeat exactly | Errors not fed back / local optimum | Action-hash dedup + breaker |
| Thrashing | Observations stagnate, trajectory oscillates | No planning, no memory | Stagnation detection + forced reflection |
| Premature exit | Early end_turn | Goal dilution / low completion bar | finish tool + environment verification gate |
A plain but life-saving practice
Before shipping any agent, run a "failure drill": deliberately give it an unsolvable task (e.g. fix a bug that doesn't exist) and watch the loop's behavior. It should concede gracefully within budget and explain where it's stuck — not burn through the limit. Behavior on unsolvable tasks predicts the severity of production incidents better than success rates on solvable ones.
References
- ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629) — the original ReAct paper, Yao et al., ICLR 2023; the basis for Section 5.
- ReAct project page — the paper's accompanying code and demos.
- Building effective agents — Anthropic Engineering — the workflow-vs-agent divide and the minimal-loop-first philosophy; cited in Sections 1 and 6.
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (arXiv:2405.15793) — the ACI design and ablations; source of the 12.5% vs 3.8% numbers; cited in Section 6.
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents (arXiv:2407.16741) — the event-stream architecture and CodeActAgent; cited in Section 6.
- Basic agentic loop with Claude and tool calling — Temporal AI Cookbook — a runnable agentic loop on the Anthropic API, cross-referencing Section 3's pseudocode.
- ollama issue #17617: model stuck in a self-sustaining tool-call loop — the real incident log of 193 identical calls and ~31M tokens; cited in Sections 4 and 7.