Skip to content

Agent Systems, Dissected: The Full Picture

At a glance Dissecting a production-grade agent system into six layers — perception, cognitive core, planning, action, memory, and governance — covering each layer's responsibilities, key design decisions, and common implementations, with a complete layered architecture diagram and a guided tour of one request through the system.

Agent Systems, Dissected: The Full Picture ​

When you're reading papers or running demos, an agent looks like just a loop: the LLM thinks, calls a tool, thinks again. But the moment you try to make it a production system — monitorable, cost-controlled, debuggable when it fails, with permissions locked down — you find the real engineering lives outside the loop: how input comes in, how context is assembled, how memory is stored and fetched, how every tool call gets audited.

This page does one thing: dissect a production-grade agent system layer by layer. For each layer we answer three questions — what it's responsible for, what the key design decisions are, and how the industry usually implements it. If you haven't yet built a basic intuition for agents, read What Is an AI Agent first and come back.

A prior judgment

The core trade-off Anthropic emphasized in Building Effective Agents (2024-12) hasn't aged: if a well-defined workflow (code orchestrating LLMs through a fixed flow) can solve the problem, don't reach for an autonomous agent (an LLM deciding its own next step in a loop). The "six-layer dissection" below describes the complete shape, but most production systems are trimmed subsets — knowing which layers you need and which are over-engineering is itself an architectural skill.

1. Overview: the Six-Layer Architecture ​

Start with the whole picture. The diagram below splits a production-grade agent system into five vertical layers plus one cross-cutting governance layer, with arrows marking the data flow of one complete task from input to output:

                        ┌─────────────────────────────────────────────┐
                        │    Governance Layer (cross-cuts all)        │
                        │  Permissions/approvals · Evals · Trace      │
                        │  observability · Cost/rate limits           │
                        └──────┬──────┬──────┬──────┬──────┬──────────┘
                               ▼      ▼      ▼      ▼      ▼
 User/event ──(1)──▶ ┌──────────────┐
 text/files/screens  │  Perception  │  Input parsing, multimodal
 /tool results       │    Layer     │  to text, session state,
                     └──────┬───────┘  input validation
                            │ (2) Structured observation
                            ▼
                     ┌──────────────┐   (3) Assemble prompt    ┌──────────────┐
                     │  Cognitive   │ ◀─────────────────────── │    Memory    │
                     │  Core        │ ─────────▶ LLM inference▶│    Layer     │
                     │  LLM + prompt│                          │ Short-term   │
                     │  + context   │                          │ Long-term/RAG│
                     │    assembly  │                          │              │
                     └──────┬───────┘                          └──────▲───────┘
                            │ (4) Output: a reply or a tool-call intent │ (7) Write
                            ▼                                           │ summaries/
                     ┌──────────────┐                                   │ facts/
                     │   Planning   │  Task breakdown, todo lists,      │ lessons
                     │    Layer     │  reflection & replanning          │
                     └──────┬───────┘                                   │
                            │ (5) Next action                           │
                            ▼                                           │
                     ┌──────────────┐                                   │
                     │    Action    │  Tool calls, code-exec sandbox,   │
                     │    Layer     │  browser/computer control         │
                     └──────┬───────┘                                   │
                            │ (6) Tool results (observation) ───────────┘
                            ▼
                    Back to the cognitive core; loop until a final
                    answer is produced ──(8)──▶ User

The loop of one task is (2)→(3)→(4)→(5)→(6)→back to (3), interleaved with memory reads (at step 3) and writes (at step 7). This is what the classic Agent Loop looks like through a layered lens: the loop itself is just the round trip between the cognitive core and the action layer; the other four layers exist to make that loop run reliably in the real world.

2. The Perception Layer: the System's Five Senses ​

Responsibility: turn signals from the external world into structured input the cognitive core can digest.

What the perception layer handles goes far beyond "the sentence the user typed":

Input typeTypical sourcesWhat the perception layer must do
User textChat box, API requests, IM botsIntent detection, session attribution, sensitive-content screening
FilesUser-uploaded PDFs / repos / spreadsheetsParsing, chunking, OCR; convert to text or structured data
Screenshots/imagesGUI screenshots, pasted imagesUnderstand directly with a multimodal model, or convert to a text description first
Tool resultsOutput of the previous tool callCleaning, truncation, error normalization
System eventsWebhooks, cron jobs, message queuesDeserialize into task descriptions with triggering context attached

Key design decisions:

  • Multimodal or textualized? Let a multimodal model look at the screenshot directly, or first have a vision model convert the screenshot into a text description for the main model? The former is higher fidelity but token-expensive; the latter is cheap but lossy. A common production pattern is tiering: main reasoning on text, native multimodality only when precise visual grounding is needed (e.g. coordinate clicks in GUI operation).
  • Tool results count as perception too. Many tutorials just concatenate tool output back into the context, but tool returns are the agent's most important perception channel. An API returning 5 MB of JSON will blow up the context window; the perception layer needs to truncate, summarize, and normalize errors into structured form ({"error": "rate_limited", "retry_after": 30} is far more useful than a wall of stack trace).
  • Input is an attack surface. Files and web pages can carry prompt injections; the perception layer is the first checkpoint for sanitization — see Security for a deeper discussion.

Common implementations: a lightweight system is just a FastAPI/Express entry point plus file-parsing libraries; heavier systems add a standalone ingestion pipeline (queue + parsing workers) that decouples "request received" from "inference started."

3. The Cognitive Core: LLM, Prompts, and Context Assembly ​

Responsibility: in every turn of the loop, based on everything currently known, decide "what to say" or "what to do."

The cognitive core isn't a single model call; it's the combination of three parts:

┌─ System prompt ── persona, rules, tool descriptions (relatively static)
│
├─ Context assembly ── pull this turn's materials from each data source:
│   conversation history + retrieved knowledge + long-term memory
│   + tool results + planning state
│
└─ LLM inference ── within a finite context window, produce:
    a final answer or a tool call

Key design decisions:

  • Context engineering is this layer's main battlefield. Anthropic's September 2025 piece Effective Context Engineering puts it bluntly: the context window is a finite, decaying resource (the so-called context rot), and the engineering goal is to give the model "just enough" information each turn — more is not better. Concrete tactics: truncating tool results on arrival, compressing conversation history (compaction), keeping bulky material in the filesystem for on-demand reading instead of stuffing it into the context. See Context Engineering for depth.
  • Layered prompts. The system prompt governs "who you are and what rules you follow"; the task prompt governs "what to do this time"; few-shot examples govern "in what format." Mixing all three into one giant prompt is the most common maintenance disaster. See Prompt Engineering.
  • Structured output. Whether the action layer can execute reliably depends on the cognitive core's tool calls strictly matching the schema. Use native tool calling / JSON-schema-constrained decoding; don't parse free text with regex.

Common implementations: one chat.completions / messages call wrapped in a context assembler. At the framework level, LangGraph models it explicitly as nodes; the OpenAI Agents SDK wraps it in a Runner — essentially the same thing. See the framework guide pages.

The cognitive core is the system's only irreplaceable layer

Perception can be reduced to a plain text entry point, planning can be dropped, memory can be nothing but in-session context, action can be one or two tools — but not the cognitive core. The flip side: when you upgrade models, this layer changes the most. Complex prompt scaffolding designed for the GPT-4 era is often a liability on new-generation reasoning models. Keeping prompts and assembly logic swappable is insurance for the future.

4. The Planning Layer: Task Breakdown, Todos, and Reflection ​

Responsibility: turn "the outcome the user wants" into "a list of executable steps," and correct course when execution drifts.

Planning has three engineering shapes, increasing in complexity and autonomy:

  1. No explicit planning: the agent decides only the next step each turn (ReAct style, arXiv:2210.03629). Fits short tasks with few tools. Most support and Q&A agents can stop here.
  2. An explicit todo list: produce a plan before executing (often a markdown checklist), then check items off during execution and allow revisions. The TodoWrite tool in Claude Code and similar coding agents is exactly this shape — the plan lives as tool state, both constraining the model and giving humans a visible progress bar.
  3. Reflection and replanning: evaluate results after execution and backtrack to revise the plan on failure. Reflexion (arXiv:2303.11366) is the academic prototype; the more common engineering form is the "verification node" — insert a check after critical steps (did the tests pass? was the retrieval relevant?) and return to planning if it fails.

Key design decisions:

  • Plan first, or think while doing? Producing a full plan up front (plan-then-execute) is controllable and reviewable but adapts poorly to a changing environment; improvising is flexible but drifts on long tasks. The practical consensus in coding agents is a hybrid: rough plan first, local replanning allowed during execution.
  • Where does the plan live? Keeping it "in the model's head" (re-derived each turn) is unreliable — on long tasks the model forgets the original plan. Putting it in explicit state (a todo tool, a scratchpad file, a graph node) is reliable, and it lets the governance layer audit "what it intends to do."
  • Reflection is not free. Every added self-evaluation is an extra LLM call, doubling latency and cost. Add verification nodes only to steps where failure is expensive; don't reflect globally.

Deeper discussion in Planning & Task Decomposition.

5. The Action Layer: Tools, Code Execution, and Computer Control ​

Responsibility: turn the cognitive core's intent into side effects on the real world, and bring the results back.

Ranked by "side-effect intensity," the action layer's tools form a spectrum:

Tool categoryExamplesRisk profile
Read-only retrievalSearch, database queries, file readingLow: watch for injection and data leakage
Code executionRunning Python/JS in a sandboxMedium: needs isolation and resource limits
Write operationsModifying files, API writes, database changesHigh: needs permission boundaries and auditing
Irreversible external actionsSending email, placing orders, deleting resources, transfersHighest: must have human-in-the-loop approval

Key design decisions:

  • The engineering quality of tool descriptions determines call quality. Tool names, parameter schemas, and docstrings are the model's only basis for choosing tools. "Returns the user object" loses to "Returns the user object with id/name/email; returns null (not an error) when not found." This is the core topic of the Tools & MCP page.
  • MCP is unifying the tool surface. The Model Context Protocol standardizes "exposing tools/data sources to agents" as a client-server protocol — the official analogy is a USB-C port for AI applications. Since 2025, mainstream clients — Claude, ChatGPT, VS Code, Cursor — all support MCP; prefer an MCP server when wiring tools into new systems rather than writing a bespoke adapter per client.
  • Isolate the execution environment. Code execution must run in a sandbox (containers, microVMs) with network allowlists and resource limits. Same for browser/computer control: give the agent a controlled virtual machine, not your dev machine.
  • Failure is a normal input. Network timeouts, rate limits, schema mismatches — the action layer must normalize every failure into an error return the model can understand, letting the cognitive core decide whether to retry, switch tools, or give up. Systems that throw exceptions straight out of the loop don't survive their first Monday.

For interception design involving human approval, see Human-in-the-Loop.

6. The Memory Layer: Short-Term Context, Long-Term Memory, and Knowledge Bases ​

Responsibility: give the agent information beyond a single request — remember this conversation, remember this user, know the organization's knowledge.

Memory is usually split into three tiers, mirroring the classic taxonomy of human memory (the mainstream academic taxonomy since 2023):

┌─ Short-term / working memory ── the session state inside the context window
│   Lifetime = one task; managed by trimming, summarization, compaction
│
├─ Long-term memory ── persisted across sessions, three kinds:
│   · Semantic: user profile, preferences, facts ("the user is a backend engineer")
│   · Episodic: records of past interactions ("last Wednesday's deploy failed")
│   · Procedural: learned workflows and skills ("this team's release process is…")
│
└─ Knowledge base / RAG ── the org's documents, codebase, tickets
    Strictly speaking not "memory" but external knowledge; retrieved, then injected into context

Key design decisions:

  • Who decides what gets remembered? Two routes: passive extraction (a background process pulls facts out of conversations after they end — Mem0 is the representative) versus agent-managed memory (the model itself calls "memory read/write" tools to decide what to store — MemGPT/Letta, arXiv:2310.08560, is the representative, with a three-tier core / recall / archival hierarchy analogous to an OS's memory management). The former is simple and controllable; the latter is flexible but harder to debug. The 2025–2026 trend is a fusion: the agent records proactively while a background process tidies asynchronously (Letta calls it sleep-time compute).
  • Wrong memories are worse than no memories. One incorrect user-profile fact pollutes every subsequent session. Long-term memory needs write-time confidence, provenance, and expiry mechanisms — this is fundamentally a data governance problem, not a retrieval problem.
  • The boundary between RAG and memory. A useful engineering distinction: memory is "about the interaction history," RAG is "about the world's knowledge." Their storage and retrieval stacks overlap heavily (vector stores + hybrid retrieval), but their update frequency, permission models, and error costs are entirely different — don't cram them into one table. See Memory Systems and RAG for depth.

7. The Governance Layer: the Control Plane Above Everything ​

The governance layer is not a stop on the data path; it's the control plane that sits above all layers. It answers four questions:

1. Permissions: what is it allowed to do? Every tool call passes a policy check: does this agent's identity have permission to call this API, touch this data? High-risk actions (sending email, changing production config) go through a human-approval queue. Least privilege matters more for agents than for people — a person abusing power is malicious; an agent abusing power may just be hallucinating. See Security.

2. Evaluation: is it doing a good job? Offline evals (regression on a fixed dataset) + online spot checks (sampling and scoring production traces). Iterating an agent without evals is refactoring with your eyes closed. See Evaluation and Evals in Practice.

3. Observability: what just happened? One agent execution is a call chain: user input → assembly → LLM → tool → LLM → … → answer. Every hop must record inputs and outputs, token consumption, latency, tool arguments and results — that's the trace. OpenTelemetry's GenAI semantic conventions already define standard span attributes for model calls, tool calls, and token usage, and platforms like LangSmith (full OTel support since March 2025), Langfuse (open source, self-hostable), and Arize Phoenix are built around this data model. Traces are simultaneously the raw material for evaluation and debugging — see Observability.

4. Cost: what did this turn cost? An agent's cost structure is nonlinear: the number of loop iterations is unpredictable, and one runaway retry loop can burn the price of a nice dinner. The governance layer needs: per-session/per-tenant budget caps, hard caps on loop iterations, routing by task complexity to models at different price points (simple classification on a small model, main reasoning on a large one), and prompt caching. For the detailed math, see Cost Control.

The governance layer is the easiest to cut — and the one you shouldn't cut

At demo stage, the governance layer's code volume is zero, and so is the chance of an incident — because there are no real users or data. After launch, it inverts: most production incidents are not "the model wasn't smart enough" but "permissions weren't locked down, failures left no trace, costs had no ceiling." Teams that treat governance as a v2 feature rarely get to ship a v2.

8. One Request's Journey Through the System ​

Stringing the six layers together, follow a concrete request — "look into last month's billing anomalies and draft an email to finance" — through the system:

#StageLayers involvedWhat actually happens in the system
1ReceivePerceptionAuth check, session restore; governance opens the request's trace root span
2UnderstandPerception → cognitive coreInput structured; the assembler pulls in conversation history, the user profile (semantic memory), and finance-knowledge-base retrieval results
3PlanPlanningProduce a todo: (1) pull billing data (2) find anomalies (3) draft the email; written into explicit plan state
4DecideCognitive coreLLM emits a tool call: query_billing(month=2026-07)
5Permission checkGovernancePolicy engine confirms this identity may read billing data; allowed, audit-logged
6ExecuteActionCall the billing API; the result is 200 KB — the perception layer truncates and summarizes it down to key fields
7LoopCognitive coreFrom the billing data the model spots 3 anomalies and decides to call get_invoice_detail for evidence
8ReflectPlanningTodo items (1) and (2) checked off; the verification node confirms "every anomaly is backed by a document" — passes
9High-risk interceptionGovernanceThe model wants to call send_email — the policy engine classifies it as a high-risk external action and downgrades it to "generate a draft for user confirmation" (human-in-the-loop)
10DeliverCognitive core → perceptionThe email draft + anomaly summary are generated and rendered to the user
11ConsolidateMemoryA background process extracts "user cares about billing anomalies" and "who is the finance contact" into long-term memory
12Wrap upGovernanceTrace closed: total turns, total tokens, total cost recorded; the session counts against the user's monthly quota

Note steps 5, 9, and 12: in a "bare loop" demo these three steps don't exist at all — and they are the watershed between a production system and a toy.

9. Single-Agent vs. Multi-Agent: How the Anatomy Differs ​

Run the same six-layer framework over single-agent and multi-agent systems and the shapes that grow out of it differ sharply:

DimensionSingle agentMulti-agent
Cognitive coreOne LLM loop, one main contextEach sub-agent has its own context; they communicate via messages
Planning layerPlanning and execution happen in one headA dedicated orchestrator emerges: decompose, dispatch, aggregate
Memory layerOne memory serving one principalAdds the "shared memory/blackboard" problem: who writes, who reads, how are conflicts resolved
Action layerA limited, cohesive toolsetToolsets isolated per role; permissions naturally least-privileged
Governance layerThe trace is one chainThe trace is a tree: cross-agent span correlation, cost attribution, and error propagation all get harder
Failure modesContext blowup, runaway loopsAll of the above ×N, plus message distortion, deadlocks, and unattributable blame

When is multi-agent actually the right call? LangChain's early-2026 guidance aligns with Anthropic's: start with a single agent and good tools; most tasks stop there. The genuine upgrade signals are:

  • A single agent's context can't hold all the domain knowledge the task needs (each sub-agent's independent context window is the most concrete benefit);
  • The tool count has grown to the point where the model's selection error rate visibly rises (split toolsets by role);
  • The task is naturally parallel (e.g. researching ten topics simultaneously);
  • Different stages need different models or different permission boundaries.

Conversely, role-play multi-agent setups — "one agent plays product manager, one plays engineer, one plays QA" — are over-engineering in most production scenarios: passing information between sub-agents in natural language is lossy compression, and coordination costs routinely exceed the gains from division of labor. For framework and topology details, see Multi-Agent Architectures.

Think of multi-agent as organizational design, not algorithm design

A single-agent system's bottleneck is model capability; a multi-agent system's bottleneck is interface design — how to slice the task, what format the handoff messages take, who arbitrates conflicts. That's the same class of problem as designing how an engineering team divides work. If your multi-agent topology looks like a real company's bloated reporting chain, it will probably be just as inefficient.

10. Summary: How to Use the Layered View ​

This dissection has three practical uses:

  1. As a design checklist. When a new system starts, walk the layers one by one: do I need this layer? In what form? The answer for most internal tools: perception kept minimal, planning omitted, memory short-term only, three to five tools, governance limited to traces and an iteration cap.
  2. As a debugging map. When an agent misbehaves, locate the layer first: irrelevant answers → look at the cognitive core's context assembly; wrong tool calls → look at the action layer's tool descriptions; duplicated work → look at the memory layer; runaway spend → look at the governance layer's caps.
  3. As a study index. This site's Core Components and Advanced Architecture sections are basically organized along these six layers — start from the learning paths and fill in whichever layer you're weakest on.

References ​