Skip to content

Case Study: Codex

At a glance Codex is not "a model that answers directly" — it's a closed-loop system driven by a harness. This article breaks down, step by step, what happens and why at every stage of its interaction, drawing the clearest possible responsibility boundary between the harness and the LLM.

Case Study: Codex ​

Codex is OpenAI's coding agent (the CLI, the desktop app, and cloud tasks all share one interaction model). Unlike the other case studies, this article doesn't rely on reverse engineering or public benchmarks. Instead it lays out a stable responsibility model: at every step of Codex's interaction with the user, what is the harness doing, on what principle, and producing what?

Why read this one

In Model vs. Harness we keep drawing a line between the "model" and the "harness." Codex is the case that makes that line sharpest: every one of its steps splits cleanly into "the harness's job" and "the LLM's job," and each half leaves behind an inspectable artifact. By the end, you'll have a checklist you can use to audit whether any agent's interaction loop actually closes.

The One-Sentence Summary ​

Codex is not "a model that receives a question and answers it directly." It's a closed-loop system driven by a harness: the harness organizes context, calls the model, executes tools, receives observations, manages permissions and state, and decides when to continue or stop; the LLM, on each turn, interprets the current state and either proposes a next action or produces a reply.

You can compress it into one line:

Codex = Harness (context engineering + tool system + control loop + state management + permission guardrails + observability) + LLM

The LLM here is the reasoning-and-generation component the harness calls; it is not the harness itself.

The Interaction at a Glance ​

                  ┌────────────┐
                  │ User input │
                  └─────┬──────┘
                        ▼
        ┌─────────────────────────────┐
        │ Parse input & triage        │
        │ instructions                │
        └───────────────┬─────────────┘
                        ▼
        ┌─────────────────────────────┐
        │      Assemble context       │
        └───────────────┬─────────────┘
                        ▼
        ┌─────────────────────────────┐
        │   Scope the task & plan     │
        └───────────────┬─────────────┘
                        ▼
      ┌──►┌─────────────────────────────┐
      │   │ Call the LLM to decide      │
      │   │ the next step               │
      │   └─────────────┬───────────────┘
      │                 ▼
      │          ┌───────────────┐
      │          │ What's next?  │
      │          └──┬───┬───┬────┘
      │             │   │   │
      │   needs a   │   │   │ needs a user
      │   tool      │   │   │ decision
      │             ▼   ▼   └──────► request confirmation
      │   ┌─────────────────┐        or input ──► back to user input
      │   │ Permission &    │
      │   │ argument checks │
      │   └────────┬────────┘
      │            ▼
      │  ┌───────────────┐      ┌───────────────────┐
      │  │ Execute tool  │─────►│ Return            │
      │  └───────────────┘      │ observation       │
      │                         └────────┬──────────┘
      │                                  ▼
      │                ┌──────────────────────────┐
      └────────────────┤ Update state & context   │◄─────────┐
                       └──────────────────────────┘          │
                                                             │
   Ready to deliver ──► verify & judge ──┬─ not met ─────────┘
                                         └─ met ──► generate final answer
                                                    ──► log & persist state

This is not a fixed pipeline that runs exactly once. The core is the Agent Loop — decide → act → observe → decide again — until the task is done, a user decision is needed, an unrecoverable error occurs, or a run boundary is hit.

What Each Step Does, and What It Produces ​

StepWhat the harness is doingCore principlePrimary output
1. Receive inputTake in text, images, files, code references, and environment infoMultimodal input is first treated as raw material to be parsed; "what the material says" and "what the user wants done" must be kept separateThe raw request object, attachment references, session metadata
2. Separate instructions from dataRecognize system rules, developer rules, project rules, the current user request, and ordinary content inside attachmentsInstructions have priority and scope; commands quoted inside referenced documents are data by default and never automatically gain the authority of user instructionsThe effective instruction set, conflicts, constraint list, open questions
3. Assemble contextSelect what the model actually needs this turnThe model can only work from the context in front of it; the harness trades information completeness against context cost through retrieval, trimming, summarization, and injectionThis turn's context package: the request, rules, history, relevant files, tool descriptions, environment state
4. Scope the taskNail down the goal, scope, success criteria, input evidence, and uncertaintiesMaterialize "what to do" and "what counts as done" up front, to prevent goal drift during executionScope, Rubric, Intake, Uncertainty, or a lightweight task list
5. Pick the next stepCall the LLM and let it decide, based on the current state, whether to answer, read, search, edit, execute, or askWhat the LLM produces is a candidate decision, not something that has already happened in the real worldA text draft, a structured tool call, a plan update, or a clarification request
6. Validate the tool callCheck that the tool exists, the arguments are valid, and the target is within allowed boundsThe model "wanting to call" and the system "allowing execution" are two different things; the harness builds a deterministic boundary between themA validated call, an argument error, a rejection reason, or an approval request
7. Control permissions and riskDecide whether to let the action through based on the sandbox, approval policy, allowed paths, and operational riskLeast privilege and explicit authorization bound side effects; high-risk or out-of-bounds actions must never run on the model's intent aloneAllow, block, degraded execution, or an authorization prompt shown to the user
8. Execute the actionRead files, run commands, modify code, hit services, or invoke other tools in the real environmentTools are what produce facts and side effects; only a successful tool return means the action actually happenedstdout, stderr, exit codes, file changes, API responses, screenshots, and so on
9. Form the observationTurn tool results into something the model can consume on its next turnExecution results are external observations, not the model's original prediction; surprises overturn old assumptionsAn observation: success signals, errors, diffs, logs, test results, and new facts
10. Update state and re-planWrite the observation back into task state and judge whether the original plan still holdsThe Agent Loop is not a mechanical retry; if new evidence invalidates the plan, revise the task model before acting againThe updated plan, completed items, remaining items, risk items, and the basis for the next decision
11. Manage context and memoryCompress long histories, save the necessary task state, and restore persisted information on demandThe context window is finite; keeping decisions, evidence, and outstanding obligations matters more than keeping every original wordCompressed summaries, task state, resumable notes, persisted files, or memory entries
12. Verify the resultsRun tests, builds, type checks, linters, screenshot checks, or inspect the artifacts directly"The code is written" is not "the goal is met"; completion must be backed by evidence matched to the success criteriaVerification commands and results, acceptance evidence, failures, remaining risks
13. Judge the stop conditionDecide whether to keep looping, wait for the user, report a blocker, or finish the taskStopping is decided by completion criteria and run boundaries, not by how the model feelsA state of continue, await input, blocked, failed, or complete
14. Produce the final answerShape the results into an actionable conclusion for the user, rather than dumping the internal processOutput should follow a decision structure: conclusion first, then changes, evidence, and risks; there is no need to expose a verbatim private chain of thoughtThe final answer, file links, a change summary, verification results, open items
15. Log and observeSave the necessary call records, state changes, timings, errors, and artifact indexesObservability is what makes failures diagnosable, behavior auditable, and tasks resumableStructured logs, audit events, run metrics, session or task state

What Happens Inside One Pass of the Agent Loop ​

Every pass can be compressed into five inspectable stages:

StageQuestionOutput
Read the stateWhat do we actually know right now, and what is still guesswork?Current facts and uncertainties
Choose an actionWhat's the smallest next step that either shrinks uncertainty or moves the goal forward?One answer, one tool call, or one question
Check the boundariesIs the action permitted, authorized, reversible, and unambiguous in its arguments?Allow, block, or approval
Observe the outside worldWhat did the action actually produce?Tool results and side effects
Transition the stateDoes the evidence meet the bar, and does the plan need revising?The next pass's state, or a terminal state

So the LLM's job is closer to "propose the next action based on the state," while the harness's job is "make that action actually happen in a controlled environment, and feed the result back into the loop."

Three Kinds of Output ​

1. User-Visible Output ​

  • Progress state: what it's reading right now, why, and what it expects to verify.
  • Decision requests: high-risk actions or scope-changing trade-offs that need the user's approval.
  • Final deliverables: answers, code, docs, screenshots, reports, links, and verification conclusions.

2. Internal Structured Output ​

  • Effective instructions and constraints.
  • The plan, step states, and stop conditions.
  • Tool names, arguments, and call results.
  • Permission decisions, error types, retry state, and context summaries.

3. Persistent Output in the Environment ​

  • Files created or modified.
  • Git working-tree diffs, build artifacts, and test reports.
  • Records or state changes in external services.
  • Resumable task state, logs, or memory.

A model-generated "I've finished the changes" is just text; only when a file diff and verification results exist does it become a verifiable, real output.

When the User Sends a Message Mid-Run ​

The user can append messages while the Agent Loop is running. The harness treats the new message as a fresh, high-priority session input and classifies it as one of:

  • A goal replacement: stop or abandon the old path and re-scope around the new goal.
  • An added constraint: keep the original task and update the scope or acceptance criteria.
  • A status query: report current milestones first, then resume the original task.
  • An approval reply: resume the corresponding branch based on the user's allow or deny.

The output of this step is not simply "append one more line to the chat transcript" — it's an updated task state, plan, and next action.

Failures, Exceptions, and Recovery ​

SituationHow to handle itWhat it should produce
Tool errorRead the error facts first, then judge whether it's an argument, environment, or implementation problemAn error classification, a corrective action, and a preserved failure log when needed
Observation contradicts the hypothesisTreat the anomaly as evidence against the old approach; don't keep pretending the plan holdsUpdated assumptions, a revised plan, a narrowed verification scope
Insufficient permissionDon't bypass the guardrails; request authorization or find an in-bounds alternativeAn approval request, a constrained explanation, or a safe fallback path
Context overflowCompress the history, but keep the goal, constraints, evidence, and outstanding obligationsA resumable summary and continuous task state
Cannot verifyDon't claim completion; state clearly what went unverified and what its risks areAn in progress or blocked state, or a delivery flagged with remaining risks
Unclear user intentDo low-risk observation first; ask the user only when a choice would change cost, scope, or side effectsEstablished facts and one necessary decision question

The Boundary Between Harness and LLM ​

CapabilityPrimary ownerNotes
Understanding semantics, generating candidate solutionsLLMProbabilistic judgment based on the context it's given
Selecting and trimming contextHarnessDecides what the model gets to see this turn
Tool registration and real executionHarness / tool runtimeConverts structured intent into real operations
Permissions, sandbox, approvalsHarnessEnforces deterministic constraints before an action happens
Driving the loop and stoppingHarnessContinues or ends based on tool results, task state, and boundaries
State, memory, compactionHarnessMaintains continuity across turns
Logging, tracing, auditingHarnessRecords observable events and system outcomes
Final wording of contentLLM, delivered by the harnessBounded by instructions, context, and verification evidence

The most important boundary: the LLM can propose an action and interpret a result, but whether files actually changed, whether a command succeeded, whether permissions allow it, and whether tests pass — all of that is decided by evidence returned by the harness and the tools.

Mapping It to a Real Bug Fix ​

  1. The user submits the bug symptoms and the repository.
  2. The harness parses the request and loads project rules and available tools.
  3. Codex pins down scope and success criteria — for example: "reproduce the failure, fix the root cause, relevant tests pass."
  4. The LLM decides to read the relevant files and tests first.
  5. The harness validates the read arguments, executes, and returns the real file contents.
  6. Based on the observations, the LLM proposes a minimal fix and issues an edit call.
  7. The harness applies the changes and returns the file diff to the model.
  8. The LLM asks to run the targeted tests; the harness executes them and returns exit codes and logs.
  9. If tests fail, the error becomes a new observation and the loop returns to diagnosis; if they pass, it reviews the diff and the broader blast radius.
  10. Once the acceptance criteria are met, Codex reports what changed, the verification evidence, and the remaining risks.

The delivery chain you can actually trust looks like this:

User goal → acceptance criteria → file changes → execution results → verification evidence → final answer

If any link is missing, the loop should be treated as not yet closed.

What to Learn from Codex ​

  1. A step only counts as having happened if it produces an artifact. Each of Codex's 15 steps defines a concrete output — the context package, the validated call, the observation, the acceptance evidence. When you audit your own harness, any step that can't name its artifact is probably just "the model's feeling" rather than a controlled system behavior.
  2. There must be a deterministic boundary between "wants to call" and "is allowed to execute." An LLM's decisions are probabilistic candidate actions; tool validation and permission gates are deterministic checks. Merge the two layers, and you're letting the model's hallucinations write directly into the real world.
  3. Stop conditions belong to the system, not the model. continue / await input / blocked / complete are decided by completion criteria and run boundaries. An agent that stops whenever it "feels done" will, on the hardest tasks, stop precisely one step short of delivery.
  4. Every link in the delivery chain must be backed by evidence. On the chain "user goal → acceptance criteria → file changes → execution results → verification evidence → final answer," each link can only be supported by the evidence of the link after it — never by the claims of the link before it.

Scope and Caveats ​

This article describes a stable responsibility model for Codex-style agents — the interaction logic shared by the CLI, the desktop app, and cloud tasks. Individual versions may differ in tool names, permission policies, context compaction algorithms, memory scope, and logging implementation.

"Thinking" here means explainable task judgment — plans, hypotheses, rationale, next steps — not a published, verbatim transcript of the model's private chain of thought. What's genuinely valuable and verifiable to the user is the decision, the action taken, the external observation, and the verification evidence.

Further Reading ​