Appearance
Case Study: Codex
Codex is OpenAI's coding agent (the CLI, the desktop app, and cloud tasks all share one interaction model). Unlike the other case studies, this article doesn't rely on reverse engineering or public benchmarks. Instead it lays out a stable responsibility model: at every step of Codex's interaction with the user, what is the harness doing, on what principle, and producing what?
Why read this one
In Model vs. Harness we keep drawing a line between the "model" and the "harness." Codex is the case that makes that line sharpest: every one of its steps splits cleanly into "the harness's job" and "the LLM's job," and each half leaves behind an inspectable artifact. By the end, you'll have a checklist you can use to audit whether any agent's interaction loop actually closes.
The One-Sentence Summary
Codex is not "a model that receives a question and answers it directly." It's a closed-loop system driven by a harness: the harness organizes context, calls the model, executes tools, receives observations, manages permissions and state, and decides when to continue or stop; the LLM, on each turn, interprets the current state and either proposes a next action or produces a reply.
You can compress it into one line:
Codex = Harness (context engineering + tool system + control loop + state management + permission guardrails + observability) + LLM
The LLM here is the reasoning-and-generation component the harness calls; it is not the harness itself.
The Interaction at a Glance
┌────────────┐
│ User input │
└─────┬──────┘
▼
┌─────────────────────────────┐
│ Parse input & triage │
│ instructions │
└───────────────┬─────────────┘
▼
┌─────────────────────────────┐
│ Assemble context │
└───────────────┬─────────────┘
▼
┌─────────────────────────────┐
│ Scope the task & plan │
└───────────────┬─────────────┘
▼
┌──►┌─────────────────────────────┐
│ │ Call the LLM to decide │
│ │ the next step │
│ └─────────────┬───────────────┘
│ ▼
│ ┌───────────────┐
│ │ What's next? │
│ └──┬───┬───┬────┘
│ │ │ │
│ needs a │ │ │ needs a user
│ tool │ │ │ decision
│ ▼ ▼ └──────► request confirmation
│ ┌─────────────────┐ or input ──► back to user input
│ │ Permission & │
│ │ argument checks │
│ └────────┬────────┘
│ ▼
│ ┌───────────────┐ ┌───────────────────┐
│ │ Execute tool │─────►│ Return │
│ └───────────────┘ │ observation │
│ └────────┬──────────┘
│ ▼
│ ┌──────────────────────────┐
└────────────────┤ Update state & context │◄─────────┐
└──────────────────────────┘ │
│
Ready to deliver ──► verify & judge ──┬─ not met ─────────┘
└─ met ──► generate final answer
──► log & persist stateThis is not a fixed pipeline that runs exactly once. The core is the Agent Loop — decide → act → observe → decide again — until the task is done, a user decision is needed, an unrecoverable error occurs, or a run boundary is hit.
What Each Step Does, and What It Produces
| Step | What the harness is doing | Core principle | Primary output |
|---|---|---|---|
| 1. Receive input | Take in text, images, files, code references, and environment info | Multimodal input is first treated as raw material to be parsed; "what the material says" and "what the user wants done" must be kept separate | The raw request object, attachment references, session metadata |
| 2. Separate instructions from data | Recognize system rules, developer rules, project rules, the current user request, and ordinary content inside attachments | Instructions have priority and scope; commands quoted inside referenced documents are data by default and never automatically gain the authority of user instructions | The effective instruction set, conflicts, constraint list, open questions |
| 3. Assemble context | Select what the model actually needs this turn | The model can only work from the context in front of it; the harness trades information completeness against context cost through retrieval, trimming, summarization, and injection | This turn's context package: the request, rules, history, relevant files, tool descriptions, environment state |
| 4. Scope the task | Nail down the goal, scope, success criteria, input evidence, and uncertainties | Materialize "what to do" and "what counts as done" up front, to prevent goal drift during execution | Scope, Rubric, Intake, Uncertainty, or a lightweight task list |
| 5. Pick the next step | Call the LLM and let it decide, based on the current state, whether to answer, read, search, edit, execute, or ask | What the LLM produces is a candidate decision, not something that has already happened in the real world | A text draft, a structured tool call, a plan update, or a clarification request |
| 6. Validate the tool call | Check that the tool exists, the arguments are valid, and the target is within allowed bounds | The model "wanting to call" and the system "allowing execution" are two different things; the harness builds a deterministic boundary between them | A validated call, an argument error, a rejection reason, or an approval request |
| 7. Control permissions and risk | Decide whether to let the action through based on the sandbox, approval policy, allowed paths, and operational risk | Least privilege and explicit authorization bound side effects; high-risk or out-of-bounds actions must never run on the model's intent alone | Allow, block, degraded execution, or an authorization prompt shown to the user |
| 8. Execute the action | Read files, run commands, modify code, hit services, or invoke other tools in the real environment | Tools are what produce facts and side effects; only a successful tool return means the action actually happened | stdout, stderr, exit codes, file changes, API responses, screenshots, and so on |
| 9. Form the observation | Turn tool results into something the model can consume on its next turn | Execution results are external observations, not the model's original prediction; surprises overturn old assumptions | An observation: success signals, errors, diffs, logs, test results, and new facts |
| 10. Update state and re-plan | Write the observation back into task state and judge whether the original plan still holds | The Agent Loop is not a mechanical retry; if new evidence invalidates the plan, revise the task model before acting again | The updated plan, completed items, remaining items, risk items, and the basis for the next decision |
| 11. Manage context and memory | Compress long histories, save the necessary task state, and restore persisted information on demand | The context window is finite; keeping decisions, evidence, and outstanding obligations matters more than keeping every original word | Compressed summaries, task state, resumable notes, persisted files, or memory entries |
| 12. Verify the results | Run tests, builds, type checks, linters, screenshot checks, or inspect the artifacts directly | "The code is written" is not "the goal is met"; completion must be backed by evidence matched to the success criteria | Verification commands and results, acceptance evidence, failures, remaining risks |
| 13. Judge the stop condition | Decide whether to keep looping, wait for the user, report a blocker, or finish the task | Stopping is decided by completion criteria and run boundaries, not by how the model feels | A state of continue, await input, blocked, failed, or complete |
| 14. Produce the final answer | Shape the results into an actionable conclusion for the user, rather than dumping the internal process | Output should follow a decision structure: conclusion first, then changes, evidence, and risks; there is no need to expose a verbatim private chain of thought | The final answer, file links, a change summary, verification results, open items |
| 15. Log and observe | Save the necessary call records, state changes, timings, errors, and artifact indexes | Observability is what makes failures diagnosable, behavior auditable, and tasks resumable | Structured logs, audit events, run metrics, session or task state |
What Happens Inside One Pass of the Agent Loop
Every pass can be compressed into five inspectable stages:
| Stage | Question | Output |
|---|---|---|
| Read the state | What do we actually know right now, and what is still guesswork? | Current facts and uncertainties |
| Choose an action | What's the smallest next step that either shrinks uncertainty or moves the goal forward? | One answer, one tool call, or one question |
| Check the boundaries | Is the action permitted, authorized, reversible, and unambiguous in its arguments? | Allow, block, or approval |
| Observe the outside world | What did the action actually produce? | Tool results and side effects |
| Transition the state | Does the evidence meet the bar, and does the plan need revising? | The next pass's state, or a terminal state |
So the LLM's job is closer to "propose the next action based on the state," while the harness's job is "make that action actually happen in a controlled environment, and feed the result back into the loop."
Three Kinds of Output
1. User-Visible Output
- Progress state: what it's reading right now, why, and what it expects to verify.
- Decision requests: high-risk actions or scope-changing trade-offs that need the user's approval.
- Final deliverables: answers, code, docs, screenshots, reports, links, and verification conclusions.
2. Internal Structured Output
- Effective instructions and constraints.
- The plan, step states, and stop conditions.
- Tool names, arguments, and call results.
- Permission decisions, error types, retry state, and context summaries.
3. Persistent Output in the Environment
- Files created or modified.
- Git working-tree diffs, build artifacts, and test reports.
- Records or state changes in external services.
- Resumable task state, logs, or memory.
A model-generated "I've finished the changes" is just text; only when a file diff and verification results exist does it become a verifiable, real output.
When the User Sends a Message Mid-Run
The user can append messages while the Agent Loop is running. The harness treats the new message as a fresh, high-priority session input and classifies it as one of:
- A goal replacement: stop or abandon the old path and re-scope around the new goal.
- An added constraint: keep the original task and update the scope or acceptance criteria.
- A status query: report current milestones first, then resume the original task.
- An approval reply: resume the corresponding branch based on the user's allow or deny.
The output of this step is not simply "append one more line to the chat transcript" — it's an updated task state, plan, and next action.
Failures, Exceptions, and Recovery
| Situation | How to handle it | What it should produce |
|---|---|---|
| Tool error | Read the error facts first, then judge whether it's an argument, environment, or implementation problem | An error classification, a corrective action, and a preserved failure log when needed |
| Observation contradicts the hypothesis | Treat the anomaly as evidence against the old approach; don't keep pretending the plan holds | Updated assumptions, a revised plan, a narrowed verification scope |
| Insufficient permission | Don't bypass the guardrails; request authorization or find an in-bounds alternative | An approval request, a constrained explanation, or a safe fallback path |
| Context overflow | Compress the history, but keep the goal, constraints, evidence, and outstanding obligations | A resumable summary and continuous task state |
| Cannot verify | Don't claim completion; state clearly what went unverified and what its risks are | An in progress or blocked state, or a delivery flagged with remaining risks |
| Unclear user intent | Do low-risk observation first; ask the user only when a choice would change cost, scope, or side effects | Established facts and one necessary decision question |
The Boundary Between Harness and LLM
| Capability | Primary owner | Notes |
|---|---|---|
| Understanding semantics, generating candidate solutions | LLM | Probabilistic judgment based on the context it's given |
| Selecting and trimming context | Harness | Decides what the model gets to see this turn |
| Tool registration and real execution | Harness / tool runtime | Converts structured intent into real operations |
| Permissions, sandbox, approvals | Harness | Enforces deterministic constraints before an action happens |
| Driving the loop and stopping | Harness | Continues or ends based on tool results, task state, and boundaries |
| State, memory, compaction | Harness | Maintains continuity across turns |
| Logging, tracing, auditing | Harness | Records observable events and system outcomes |
| Final wording of content | LLM, delivered by the harness | Bounded by instructions, context, and verification evidence |
The most important boundary: the LLM can propose an action and interpret a result, but whether files actually changed, whether a command succeeded, whether permissions allow it, and whether tests pass — all of that is decided by evidence returned by the harness and the tools.
Mapping It to a Real Bug Fix
- The user submits the bug symptoms and the repository.
- The harness parses the request and loads project rules and available tools.
- Codex pins down scope and success criteria — for example: "reproduce the failure, fix the root cause, relevant tests pass."
- The LLM decides to read the relevant files and tests first.
- The harness validates the read arguments, executes, and returns the real file contents.
- Based on the observations, the LLM proposes a minimal fix and issues an edit call.
- The harness applies the changes and returns the file diff to the model.
- The LLM asks to run the targeted tests; the harness executes them and returns exit codes and logs.
- If tests fail, the error becomes a new observation and the loop returns to diagnosis; if they pass, it reviews the diff and the broader blast radius.
- Once the acceptance criteria are met, Codex reports what changed, the verification evidence, and the remaining risks.
The delivery chain you can actually trust looks like this:
User goal → acceptance criteria → file changes → execution results → verification evidence → final answer
If any link is missing, the loop should be treated as not yet closed.
What to Learn from Codex
- A step only counts as having happened if it produces an artifact. Each of Codex's 15 steps defines a concrete output — the context package, the validated call, the observation, the acceptance evidence. When you audit your own harness, any step that can't name its artifact is probably just "the model's feeling" rather than a controlled system behavior.
- There must be a deterministic boundary between "wants to call" and "is allowed to execute." An LLM's decisions are probabilistic candidate actions; tool validation and permission gates are deterministic checks. Merge the two layers, and you're letting the model's hallucinations write directly into the real world.
- Stop conditions belong to the system, not the model.
continue/await input/blocked/completeare decided by completion criteria and run boundaries. An agent that stops whenever it "feels done" will, on the hardest tasks, stop precisely one step short of delivery. - Every link in the delivery chain must be backed by evidence. On the chain "user goal → acceptance criteria → file changes → execution results → verification evidence → final answer," each link can only be supported by the evidence of the link after it — never by the claims of the link before it.
Scope and Caveats
This article describes a stable responsibility model for Codex-style agents — the interaction logic shared by the CLI, the desktop app, and cloud tasks. Individual versions may differ in tool names, permission policies, context compaction algorithms, memory scope, and logging implementation.
"Thinking" here means explainable task judgment — plans, hypotheses, rationale, next steps — not a published, verbatim transcript of the model's private chain of thought. What's genuinely valuable and verifiable to the user is the decision, the action taken, the external observation, and the verification evidence.
Further Reading
- The Agent Loop — the general form of the "decide → act → observe" loop
- Context Engineering — the full treatment behind step 3 ("assemble context") and step 11 ("manage context and memory")
- Permissions, Safety & Human-in-the-Loop — how to design the deterministic boundaries of steps 6 and 7
- Evaluation & Observability — how the verification of step 12 and the audit trail of step 15 get implemented
- Case Study: Claude Code — the same responsibility model implemented in another product: tool calls + permission gates + subagents
- Case Study: Aider — the contrasting sample that pushes "getting the edit onto disk" to its extreme