Appearance
The Agent Loop
The Concept: The Heart Is Just a While Loop
Strip away the frameworks, the protocols, and the marketing gloss, and the core control flow of an AI agent fits on a sticky note:
text
while not done:
context = assemble(history, memory, tools) # Observe: assemble everything the model will see
response = llm(context) # Think: the model produces text or a tool call
if response.has_tool_calls():
results = execute(response.tool_calls) # Act: the harness executes tools on its behalf
history.append(response, results) # Append the observations back into the context
else:
done = True # No tool call = the model thinks the job is doneThis is the agent loop: the model proposes, the harness executes and feeds the results back, and around it goes—until the task is done or the loop is forced to stop. In Building effective agents (December 2024), Anthropic defined an agent as "a system in which LLMs dynamically direct their own processes and tool usage." That "dynamically direct" is exactly what materializes as this loop in code: how far each iteration goes, and when to exit, are never hard-coded by the programmer—the model decides in the moment.
This is also the dividing line between an agent and a workflow. A workflow's control flow lives in the code (retrieve first, then summarize, then generate); an agent's control flow lives in the loop—the code only provides primitives (tools, context, termination conditions), and the model decides the orchestration.
Terminology
Academia often calls this structure the observe–think–act loop, after the agent–environment interaction paradigm from reinforcement learning. In the LLM setting: the environment = the file system, terminal, browser, and APIs; the observation = tool execution results; the thinking = the model's reasoning output (often natural-language reasoning emitted just before a tool call).
Why You Need a Loop: One Inference Can't Solve Real Tasks
A single LLM call is a stateless function: f(prompt) -> text. But real tasks have three properties that force you to wrap it in a loop:
- Information reveals itself gradually. Fixing a bug means reading the error first, then locating the file, then tracing the call chain—what you look at next depends on what you just saw. No static orchestration can enumerate that path in advance.
- Actions change the world. Editing files, running tests, and sending requests all mutate the environment, and every subsequent decision must be based on the new state. The loop is the only channel for feeding the consequences of an action back into decision-making.
- Errors are the norm, and recovery runs on feedback. Misspelled commands, failing tests, API timeouts—the model can only self-correct if it sees the failure. Appending the error output verbatim to the context is the simplest and most effective error-correction mechanism there is.
Put another way: the model supplies the judgment about what to do next; the loop supplies the persistence to carry that judgment through. Without the loop, the model is just a consultant; with it, the model becomes a system that gets work done. This is why, on benchmarks like SWE-bench, differences in results across the same models come mostly from the harness—the loop and its surrounding machinery (context assembly, tool design, termination checks) are the heart of the harness. See What Is an Agent Harness and Model vs. Skeleton.
Anatomy of an Iteration: What the Harness Does on Every Turn
"The model thinks one step" sounds simple, but each turn of the loop asks the harness to do six jobs, and every one of them involves real engineering decisions:
text
┌────────────────────────── One iteration ─────────────────────────────┐
│ │
▼ │
┌───────────────────────┐ ┌───────────────┐ ┌───────────────────────────────┐ │
│ 1. Assemble context │──▶│ 2. Call model │──▶│ 3. Parse output │ │
│ system prompt │ │ streaming │ │ validate tool call args │ │
│ history (compressible)│ │ timeout/retry │ │ no tool call → stop check │ │
│ tool schemas │ │ interruptible │ └──────────────┬────────────────┘ │
│ memory / retrieval │ └───────────────┘ │ │
└───────────────────────┘ ▼ │
┌───────────────────────────────┐ │
┌──────────────────────────────────│ 4. Permission → execute tool │──┘
│ │ timeout / sandbox / parallel │
▼ └───────────────────────────────┘
┌────────────────────────────┐ ┌────────────────────────────────────────┐
│ 6. Check stop conditions │◀──│ 5. Append observation │
│ done? over budget? │ │ truncate/summarize → history │
│ breaker? human interrupt? │ │ update cost / step counters │
└────────────────────────────┘ └────────────────────────────────────────┘1. Assemble context. At the start of every turn, the harness stitches the system prompt, conversation history, tool definitions, and memory retrieval results into one complete model call. Context doesn't only grow—when it exceeds the limit, it must be compressed or truncated, which is the central problem of Context Engineering.
2. Call the model. Usually a streaming call (see "Streaming and Interruption Handling" below) with timeouts and retries. Production systems also have to handle rate limiting (HTTP 429) and load balancing.
3. Parse the output. The model produces two kinds of output: natural language (its reasoning, explanations for the user) and structured tool calls. The harness must validate tool call arguments against each tool's JSON Schema—models produce invalid arguments all the time, and unvalidated input must never go straight into a tool.
4. Execute tools. Before execution, calls pass through the permission system (dangerous commands require user approval); during execution, they run with timeouts and output truncation. Note that the executor of a tool is the harness, not the model. The model only produces an intention, and this separation is the foundation of every security boundary.
5. Append observations. Append the tool call and its result to the history. The crucial detail is result preprocessing: a single grep can return 10MB of text, and stuffing that straight back into the context would blow the window—it has to be truncated, paginated, or summarized.
6. Check termination conditions. The most easily underestimated step—and the one that most determines how the system behaves. It deserves a section of its own.
Termination Conditions: The Art of Stopping
An agent that doesn't know when to stop is more dangerous than one that isn't very smart—it will burn through the budget, mangle files, and spin in infinite loops. Production-grade loops usually stack four layers of termination conditions:
| Termination type | Trigger | Typical failure | Engineering notes |
|---|---|---|---|
| Task completion | The model emits no tool call this turn and outputs its final answer | The model may "fake completion"—claiming a fix without verifying it | Back it up with verification tools (e.g., require tests to pass before completion counts), or add a confirmation round after the final answer |
| Step / budget limits | Any of iteration count, token consumption, dollar cost, or wall-clock time exceeds its cap | A hard cutoff leaves the work half-finished | Inject a reminder at the soft limit ("3 steps left—wrap up"), and only cut off at the hard limit |
| Consecutive-error circuit breaker | N tool failures in a row, the model repeating the same tool call, or context over the limit with no way to compress | A threshold set too tight kills legitimate long tasks | Distinguish "the same error repeating" (a danger signal) from "normal exploration across different errors" |
| Human interruption | The user hits Esc / Ctrl-C, or a permission request is denied | Inconsistent state after the interruption (files half-edited) | Deliver interruptions as collaboration signals, not exceptions—feed "the user denied this action" back to the model as an observation so it can take another route |
"The model says it's done" ≠ "the task is done"
The most common failure mode isn't an infinite loop but quitting early: the model edits the code, never runs the tests, and declares "fixed." The countermeasure is not to trust the model's word but to build verification into the task itself—for example, the harness checks whether at least one test ran after receiving the completion signal, or an external evaluation script decides completion outright (which is exactly how SWE-bench works). See Evaluation and Observability.
Error Recovery and Retries: Turning Failure into Context
Errors inside the loop fall into two classes, and they are handled completely differently:
Transient errors (retry): network timeouts, API rate limits, malformed model output (like an unclosed JSON in a tool call). These are invisible to the model—the harness retries silently with exponential backoff. For malformed output, append the parser's error itself as feedback ("your tool call arguments are invalid: Unexpected end of JSON"), and the model will usually fix itself on the next turn.
Task-level errors (feedback): the tool ran fine, but the result is a failure—a compiler error, a red test, a missing file. These should not be retried; they should be fed back into the loop verbatim. This is the most counterintuitive property of the agent loop compared with traditional programs: in conventional code, an exception means the flow is broken; in an agent loop, the failure output is the most valuable new information available. Shown AssertionError: expected 3, got 4, the model knows on its own where to go fix things.
In practice there's one more line to draw: the third time the same error appears, escalate—inject a firm nudge ("this path is a dead end, try another"), make the model explain why it's failing before continuing, or trip the circuit breaker outright. Letting the model wrestle with the same error message for twenty rounds is the fastest way to burn tokens.
Streaming and Interruption Handling
A production-grade loop has to stream, for two reasons:
- Perceived latency. A single iteration can involve tens of seconds of reasoning. Streaming shows the thinking to the user in real time—the main reason tools like Claude Code feel fast. The total wall time doesn't change, but the user can see what it's doing.
- Interruptibility. When users see the model heading the wrong way, they must be able to stop it immediately—not wait for it to finish writing 500 lines of wrong code. Implementation-wise, an interruption is three layers working together: the transport layer cancels the HTTP stream, the execution layer sends termination signals to the running tool (SIGTERM → SIGKILL), and the logic layer appends "interrupted by the user" to the history as an observation, so the model knows what happened and knows to converge.
An interruption is input, not an exception
Modeling the user's interrupt as an exception (try/except InterruptedError, then exit) is a classic beginner mistake. Model it instead as a new user message: the loop doesn't exit—the current action is discarded and control returns to the user. The user may correct course ("don't touch that file, the problem is in the config"), and the loop keeps turning with the new information. This design choice determines whether your tool is a brittle automaton or a collaborator.
From ReAct to Production Loops: A Short History
October 2022: ReAct. Shunyu Yao et al.'s paper ReAct: Synergizing Reasoning and Acting in Language Models (published at ICLR 2023) made the first systematic case for interleaving reasoning traces (Thought) with actions (Action): at each step the model writes out its reasoning, then fires an action, and once the environment's observation (Observation) comes back, the reasoning continues. In effect, this paper legislated the agent loop—nearly every agent framework since has had a control flow that's a variation on observe–think–act. Worth noting: in the ReAct era there was no tool calling API yet; an "action" was a plain text line in the model's output (Action: search[query]), parsed with regex.
2023: the AutoGPT hype and the hangover. AutoGPT packaged the loop as a "fully autonomous agent" and exposed every weakness of a naked loop: goal drift, infinite loops, unbounded spending. The lessons crystallized into two points of consensus: loops need hard budgets, and autonomy needs to be constrained by permissions and confirmation mechanisms. That same year, OpenAI shipped function calling, and the tool call went from "text parsed by regex" to "structured output with a schema," dramatically simplifying the parsing layer.
2024: engineering convergence. SWE-agent (NeurIPS 2024) demonstrated that the harness's interface design—what the paper calls the Agent-Computer Interface (ACI)—matters to SWE-bench scores as much as the model itself: which file-editing commands you give the model, and how you format error feedback, directly determine the loop's efficiency. That same year, Anthropic's Building effective agents distilled industry practice into a single thesis: the most successful agent implementations tend to be the simplest—a while loop plus a few tools, not elaborate framework orchestration. Platforms like OpenHands (ICLR 2025) turned event streams, sandboxed execution, and concurrent tool calls into standardized loop infrastructure.
The direction of evolution is clear: the loop itself never got more complicated—what got more complicated is the quality of service the harness delivers on every turn: steadier parsing, smarter context management, finer-grained permissions, better observability. Models keep getting stronger while loops keep getting plainer; those are two faces of the same trend.
A Production-Grade Loop Skeleton (About 60 Lines of Pseudocode)
Pack all of the mechanisms above into one loop and it looks roughly like this:
python
def agent_loop(task, budget=Budget(max_steps=50, max_cost_usd=2.0)):
history = [system_prompt(), user_message(task)]
consecutive_failures = 0
while True:
# ── Termination checks: budget and circuit breaker ─────────────
if budget.exhausted():
return finish("budget limit exceeded", history)
if consecutive_failures >= 3:
return finish("consecutive failures, circuit breaker tripped", history)
# ── Assemble context (compress history if necessary) ───────────
context = assemble_context(history)
if token_count(context) > CONTEXT_LIMIT:
history = compress(history) # see the chapter on context engineering
context = assemble_context(history)
# ── Call the model (streaming, interruptible, retry transient errors) ──
try:
response = call_llm_with_retry(
context, stream=True, retries=3, backoff=exponential)
except UserInterrupt as e: # user interrupt = new input, not an exception
history.append(user_message(e.feedback or "interrupted the current action"))
continue
except LLMError as e: # retries exhausted, the service is really down
return finish(f"model service unavailable: {e}", history)
# ── No tool call = the model considers the task complete ───────
if not response.tool_calls:
if not verify_completion(task, history): # verbal completion doesn't count
history.append(user_message(
"run the tests to verify the fix before declaring it done"))
continue
return finish(response.text, history)
# ── Execute tool calls one by one ──────────────────────────────
for call in response.tool_calls:
history.append(assistant_tool_call(call))
try:
validate_args(call) # schema validation
require_permission(call) # dangerous operations need approval
result = execute_tool(call, timeout=120) # sandbox + timeout
observation = truncate(result, max_chars=30_000)
consecutive_failures = 0
except ValidationError as e: # invalid arguments → feedback to the model
observation, consecutive_failures = f"invalid arguments: {e}", +1
except PermissionDenied as e: # user declined → collaboration signal
observation = f"user denied this action: {e.reason}"
except ToolTimeout:
observation, consecutive_failures = "execution timed out (120s)", +1
except ToolError as e: # task-level failure = valuable information
observation = truncate(f"tool error: {e}")
history.append(tool_result(call.id, observation))
budget.record(call, response.usage) # bookkeeping: steps, tokens, costThis skeleton deliberately omits much of what a production system needs—event sourcing (persisting every step as a replayable log), parallel tool calls, subagent spawning, telemetry reporting—but the skeleton of the loop never changes: assemble → call → parse → execute → append → check stop. To build a minimal working version yourself, see Build Your First Harness.
Trade-offs
Loop autonomy vs. controllability. The looser the termination conditions, the harder the tasks an agent can chew through—and the higher the risk of losing control and the cost. Interactive tools (Claude Code) lean toward loose limits plus human interruption; batch evaluation (SWE-bench) leans toward strict budgets, because nobody is watching. Picking the wrong setting for your scenario is more fatal than picking the wrong model.
What to expose to the model. Error output, token consumption, steps remaining—each one occupies context and shapes model behavior. Exposing error details is almost always right (the model self-corrects with them); exposing the budget needs care—some teams have found that once the model knows "3 steps left," it starts wrapping up sloppily. Information is both a tool and a distraction.
When and how to compress history. As context approaches the limit, truncation loses early key decisions, summarization introduces distortion, and doing nothing errors out outright. There's no silver bullet, only "pick a strategy by task type"—see Context Engineering for details.
Determinism vs. flexibility. You can stuff more code logic into the loop (enforced workflows, hard-coded checkpoints) and make every step more controllable, but you'll slide back into workflow territory step by step, forfeiting the core value of an agent: handling paths you didn't anticipate. Anthropic's advice runs the other way—write the simplest loop first, and add targeted constraints only when you hit specific failures.
Further Reading
- What Is an Agent Harness — where the loop sits within the harness as a whole
- Overall Architecture Anatomy — how the loop fits together with the other components
- Context Engineering — the full deep-dive into step one of every iteration: assembling context
- The Tool System — schema design for tool calls and the ACI idea
- Permissions and Security Boundaries — approval mechanisms before executing tools
- Evaluation and Observability — how to measure a loop's performance
- SWE-agent Case Study — how ACI design changes loop efficiency
- Claude Code Case Study — the benchmark implementation of an interactive loop
- Core Papers: A Guided Tour — close readings of foundational papers like ReAct
- Common Pitfalls — a checklist of anti-patterns for loops gone wrong
References
- ReAct: Synergizing Reasoning and Acting in Language Models (arXiv:2210.03629) — Yao et al., ICLR 2023
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (arXiv:2405.15793) — Yang et al., NeurIPS 2024
- Building effective agents — Anthropic engineering blog, December 2024
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents (arXiv:2407.16741) — Wang et al., ICLR 2025