Skip to content

Case Study: OpenHands

At a glance A deep dive into the harness architecture of the open-source agent platform OpenHands (formerly OpenDevin): the event stream, Docker sandbox runtime, CodeAct's unified action space, microagents and pluggable design — plus its academic value as a "harness research platform."

Case Study: OpenHands ​

If SWE-agent demonstrated how powerful a carefully polished ACI (agent-computer interface) can be, OpenHands answers a different question: can the harness itself be built as a general-purpose, reproducible experimental platform — where anyone can swap agents, models, and tools, and run benchmarks, all on the same infrastructure?

OpenHands (formerly OpenDevin) is probably the most thorough embodiment to date of "the harness as a research object." It isn't an implementation of some agent — it's a complete scaffold around agents: an event stream, a sandboxed runtime, a skill library, multi-agent delegation, an evaluation framework. Understanding it is roughly equivalent to understanding the full checklist of a modern coding agent harness.

The project at a glance ​

DimensionDetails
Former nameOpenDevin, started March 2024 as an open-source replica of Devin; renamed OpenHands in September 2024
PaperOpenHands: An Open Platform for AI Software Developers as Generalist Agents (arXiv:2407.16741, v1 2024-07-23)
Core authorsXingyao Wang, Graham Neubig, et al. (UIUC / CMU / All Hands AI, among others)
Theoretical basisThe CodeAct paper: Executable Code Actions Elicit Better LLM Agents (arXiv:2402.01030, ICML 2024)
LicenseMIT (commercial use allowed)
Community scaleAt publication: 32K GitHub stars, 188 contributors, 2.1K+ contributions; still growing since
PositioningA general agent development platform: agent abstraction + event stream + sandbox runtime + AgentHub + eval framework

A note on project evolution

OpenHands iterates fast. This page dissects its classic architecture (matching the 2407.16741 paper and the 0.x-era code structure). In November 2025 the project released V1 along with a standalone OpenHands Software Agent SDK; the repo moved from All-Hands-AI/OpenHands to OpenHands/OpenHands, and the product expanded into Agent Canvas, which can host third-party agents (Claude Code, Codex, etc.). The architectural ideas carry through, but for specific class names and directories, defer to the version of the code you're reading.

Why OpenHands is worth studying ​

Most coding agent products (Claude Code, Cursor, Aider) hide the harness inside the product: you can observe behavior and outcomes, but not the full mechanism. OpenHands does the opposite:

  • All mechanisms in the open. Event stream, state, actions, observations, runtime — every concept has a clear code entity, and is written up in the paper.
  • Controllable variables. The agent implementation, the LLM, the sandbox image, the skill library, the benchmarks — all are swappable slots, naturally suited to ablation experiments.
  • Academic backing. Two papers (the platform paper + the CodeAct paper) spell out the motivation behind design decisions — no reverse engineering required.

In the language of the anatomy page: OpenHands is the project that made every harness component a "first-class citizen." Its own positioning is blunt — a community-driven platform, not an agent.

Overall architecture: the big three ​

The platform paper condenses OpenHands into three main components: the agent abstraction (community-contributed agent implementations, gathered in AgentHub), the event stream (a history of actions and observations), and the runtime (executes actions into observations).

┌─────────────────────── OpenHands host process ───────────────────────┐
│                                                                      │
│   User / UI / CLI                                                    │
│        │  MessageAction / feedback                                   │
│        ▼                                                             │
│   ┌─────────────┐    step(state)     ┌──────────────────────┐        │
│   │   Agent     │ ◄───────────────── │        State         │        │
│   │ (CodeAct et │                    │ ├ event stream       │        │
│   │  al.)       │                    │ │  (action+obs       │        │
│   └──────┬──────┘                    │ │   history)         │        │
│          │ returns an Action         │ ├ cumulative LLM cost │        │
│          ▼                           │ └ delegation metadata │        │
│   ┌──────────────── Event Stream ────────────────┐    └──────────┘    │
│   │  Action → dispatched to Runtime →            │                    │
│   │  Observation written back to the stream      │                    │
│   └──────────────────────┬───────────────────────┘                    │
└──────────────────────────┼────────────────────────────────────────────┘
                           │ REST API (action execution requests)
                           ▼
              ┌─── Docker sandbox container (one per task session) ───┐
              │  Action Execution API Server                          │
              │   ├─ bash shell (command execution)                   │
              │   ├─ Jupyter IPython server (runs Python)             │
              │   └─ Chromium + Playwright (browser)                  │
              │  Mount: the user's workspace directory                │
              └───────────────────────────────────────────────────────┘

The most striking thing about this structure is the separation of responsibilities: the agent only "looks at state and emits actions," with no idea how actions get executed; the runtime only "executes actions and returns observations," with no idea who emitted them or why. The two are decoupled by the event stream, an append-only log. Now, block by block.

The event stream: everything is an event ​

The event stream is OpenHands' heart. Its definition (paper §2.1): a time-ordered collection of actions and observations, including both the actions the agent emits and the user's instructions and feedback.

Several key designs revolve around the stream:

State = event stream + auxiliary data. Beyond the event history, state carries the cumulative LLM cost, tracking metadata for multi-agent delegation, execution parameters, and so on. This means harness-level concerns (how much money has been spent, who delegated which subtask to whom) and the agent's "memory" are explicitly modeled in the same data structure.

Actions and observations are paired typed objects. The core action set is deliberately small and general:

ActionPurpose
CmdRunActionExecute an arbitrary bash command in the sandbox
IPythonRunCellActionExecute Python code in the sandbox's Jupyter
BrowserInteractiveActionDrive the browser via BrowserGym's domain-specific language
AgentDelegateActionDelegate a subtask to another agent
AgentFinishAction / MessageActionEnd the task / converse with a human

The corresponding observations describe environment changes: command output, Python execution results, the browser's HTML/DOM/accessibility tree/screenshots, user messages, and so on.

The event stream is both the conversation history and the audit log. Because every interaction lands on the same stream, "what context the model saw" and "what actually happened in the system" are two views of the same thing. That makes debugging, replay, and evaluation direct — the eval framework consumes exactly these trajectories to score runs.

The harness view

The event-stream architecture is essentially event sourcing applied to agent systems: don't store "current state," store "the sequence of events," and derive state by replay. The cost is serializing the history into the prompt at every step — hence the pressure on context engineering (long-session compression and truncation are handled by a dedicated condenser mechanism). See Context Engineering.

The agent abstraction: one step function ​

On top of the event stream, the interface for implementing a new agent is compressed to the extreme — the MinimalAgent skeleton in the paper's Figure 3 says it all:

python
class MinimalAgent:
    def reset(self) -> None:
        self.system_message = "You are a helpful assistant …"

    def step(self, state: State):
        # 1. Serialize the event-stream history into a message list
        messages = [{"role": "system", "content": self.system_message}]
        for prev_action, obs in state.history:
            messages.append(get_action_message(prev_action))
            messages.append(get_observation_message(obs))

        # 2. Call the LLM and parse the output into an action
        response = self.llm.do_completion(messages)
        action = self.parse_response(response)

        # 3. Return a typed action for the runtime to execute
        if self.is_finish_command(action):
            return AgentFinishAction()
        elif self.is_bash_command(action):
            return CmdRunAction(command=action.command)
        elif self.is_python_code(action):
            return IPythonRunCellAction(code=action.code)
        elif self.is_browser_action(action):
            return BrowseInteractiveAction(code=action.code)
        else:
            return MessageAction(content=action.message)

That's what the agent loop looks like in OpenHands: step(state) -> action, with the loop itself driven by the platform. The value of the abstraction: researchers only need to care about "given this history, what's the next step" — execution details, sandbox safety, cost accounting, UI rendering are all covered by the harness. That's the precondition for being an experimental platform: swapping agents costs about as much as writing a class a few dozen lines long.

CodeAct: code as the unified action space ​

OpenHands' default general agent, CodeActAgent, is based on the same team's earlier CodeAct paper (arXiv:2402.01030, ICML 2024). That paper's question: in what format should an agent's actions be expressed?

Both mainstream approaches of the time had obvious flaws:

Action formatProblem
Predefined JSON function callingThe action space is capped by the pre-declared tool catalog; one action calls one tool, with no composition
Free-text actionsFragile parsing, format drift, unreliable execution

CodeAct's answer: have the LLM emit executable Python code as the action, run it through an interpreter, observe the result, continue. Code natively carries loops, conditionals, variables, library imports, and function composition — a single code block can accomplish what JSON actions would need a dozen round trips for. Better yet, the interpreter's error messages are themselves deterministic, high-quality environment feedback the agent can use to self-correct.

The paper tested 17 LLMs on API-Bank and a newly built tool-use benchmark: CodeAct improved success rates by up to about 20% over JSON/text formats.

OpenHands engineering-izes this idea:

  • The core actions IPythonRunCellAction and CmdRunAction are CodeAct instantiated — the action space ≈ "anything runnable on a Linux machine."
  • It also stays compatible with traditional function calling: to add a new "tool," you don't touch an action schema — you write a Python function and drop it into the sandbox's IPython environment (see AgentSkills, next section).
  • The more radical corollary: the agent can build its own tools — when no ready-made API exists, it writes a Python function on the spot.

The price of a unified action space

"Everything is code" maximizes flexibility — and maximizes the safety problem: every line of model-written code gets really executed. That's exactly why OpenHands must pair with a strongly isolated sandbox runtime. CodeAct and the Docker sandbox are mutually presupposing designs on this platform, not two independent features.

The runtime sandbox: locking danger in a container ​

CodeAct lets the agent execute arbitrary code; the runtime's job is to make that safe and reproducible. Key points of OpenHands' runtime design (paper §2.2):

One container per session. Every task session starts an isolated Docker container, and every action in the event stream executes inside it. The user's working directory is mounted in with configurable options, capping the agent's blast radius at the container.

An in-container action-execution API. A REST API server runs inside the sandbox (the OpenHands Action Execution API); the host process sends it actions from the event stream, and it maintains three things and returns observations:

  1. A bash shell attached to the container's OS;
  2. A Jupyter IPython server handling interactive Python execution;
  3. A Playwright-based Chromium browser, whose action primitives come from BrowserGym (navigate, click, type, scroll), with observations including HTML, DOM, accessibility tree, screenshots, open tabs, and so on.

Arbitrary image support. The runtime has a build mechanism: take any user-provided Docker image, inject the action-execution API, and it becomes a usable sandbox. That means evaluating different projects (different languages, dependencies, OS environments) requires no harness changes — just swap images. Essential for benchmarks like SWE-bench that require reproducing an environment per repository.

Why does the browser go in the sandbox too?

OpenHands' ambition isn't "a code completion tool" but "a general digital worker": human developers consult docs, look at issues, and browse the web while working. The built-in browser lets the same agent abstraction cover web tasks like WebArena — in the paper, CodeActAgent runs software engineering, web browsing, and misc-assistant benchmarks simultaneously without any prompt changes, purely thanks to this action-space design.

AgentSkills: tools defined not by schema but by pip ​

The SWE-agent paper proved that carefully designed tools (the ACI) massively affect performance — but creating, maintaining, and distributing tools is heavy engineering. OpenHands' AgentSkills library offers a very light answer (paper §2.3):

  • A skill is just a Python package. Writing a Python function is writing a tool; it's auto-imported into the sandbox's IPython environment and the agent calls it directly via IPythonRunCellAction. No schema registration, no agent code changes.
  • Restrained inclusion criteria. The official rules admit only two kinds: things the LLM can't reliably do by writing code itself (like edit_file, which edits files by line number), and things that require calling an external model (like parse_image calling a vision model or parse_pdf for PDFs). "The LLM already knows how to read a CSV with pandas" doesn't need wrapping.
  • Maintained like real software. The skill library ships with unit tests, avoiding "the agent takes the blame, the tool stabs it in the back."

This route continues the "tools as code" discussion from the tool component page: when the action space itself is code, "adding a tool to the agent" degenerates into "installing a dependency in the environment," and the engineering complexity drops dramatically.

microagent: two confusable concepts ​

In OpenHands-speak, "micro" appears at two levels, easily conflated in the docs; worth separating:

1. Micro agent (the specialized agent in the paper). Reuses most of a general agent's implementation (like CodeActAgent), swapping only the prompt and a bit of configuration to specialize for a task. The design goal is lowering the contribution bar — you don't need to understand the whole harness; sharing a well-tuned prompt publishes a "new agent."

2. Microagents (the knowledge-fragment mechanism in the repo). This is the mechanism users actually touch: create a .openhands/microagents/ directory at the repo root, drop in Markdown files containing project-specific guidance for the agent. The official docs define two types:

TypeTriggerFrontmatterTypical use
GeneralInjected automatically every sessionOptionalrepo.md: repo structure, build commands, code conventions
Keyword-TriggeredInjected only when a specified keyword appears in the user's messageRequired (triggers field)Specialized knowledge loaded only when a particular framework/service comes up
some-repository/
└── .openhands/
    └── microagents/
        ├── repo.md            # general repo guidance (always on)
        ├── trigger_this.md    # keyword-triggered
        └── trigger_that.md    # keyword-triggered

The harness view

Microagents are essentially conditional context injection: turning "what knowledge the model should see, when" from a one-shot giant system prompt into a modular, on-demand knowledge base. The docs explicitly warn that loaded microagents occupy the context window — so keyword triggering isn't a nice-to-have, it's a necessary means of controlling prompt size. Same idea, different implementation from Claude Code's skills and layered CLAUDE.md loading.

Multi-agent delegation and the pluggable design ​

Delegation is an action, not a framework. Multi-agent collaboration in OpenHands works through AgentDelegateAction: one agent packages a subtask and delegates it to another — for example, the general CodeActAgent delegating complex web operations to the dedicated BrowsingAgent. Delegation relationships are recorded in State's metadata. The design is deliberately simple — no blackboard system, no role-orchestration engine, just "one more action in the action space." See the comparative discussion in multi-agent orchestration.

Pluggable agents: AgentHub. The platform gathers 10+ community-contributed agent implementations: the default CodeActAgent, the web-focused BrowsingAgent, GPTSwarm's optimizable-graph agent, various micro agents… They all share the same event stream and runtime, so comparisons are fair.

Pluggable LLMs: LiteLLM. Model access goes through LiteLLM as a unified gateway; switching models means changing the model string in config (e.g. openai/gpt-4o, claude-*, OpenRouter, or local vLLM/Ollama endpoints), with practical logic like 429 rate-limit retries built in. Model and harness are fully decoupled — a living specimen of this site's model vs. harness argument.

Pluggable evaluation. The platform integrates 15 benchmarks (SWE-bench, WebArena, GAIA, GPQA, HumanEvalFix, ML-Bench, BIRD, AgentBench, MINT, etc.) across the three categories of software engineering, web browsing, and misc assistant. Representative numbers from the paper: CodeActAgent v1.8 with claude-3.5-sonnet reached 26.0% on SWE-bench Lite (average cost per instance about $1.10); the same agent, without prompt changes, reached 14.5–15.3% on WebArena and about 52% on the GPQA diamond subset.

A research finding worth noting

On AgentBench, the paper observed that when switched to a weaker model (gpt-3.5-turbo), OpenHands' general agent actually did worse than the specialized baselines in the original benchmark. The authors' interpretation: general agents have a threshold requirement on the base model's instruction-following ability; below that threshold, a carefully specialized dedicated harness is more stable. Quantitative evidence for "matching the harness to the model's capability."

Trade-offs and critical observations ​

OpenHands is an excellent research platform, but every design choice has a matching cost; see them clearly before making engineering decisions:

  • The full-serialization cost of the event stream. Sending the complete history to the model at every step inflates token consumption and latency on long tasks, forcing reliance on the condenser for compression — and the compression strategy itself becomes a new research variable. For production systems, append-only event sourcing isn't necessarily the best answer to context management.
  • The tension between a general action space and a bespoke ACI. The "bash + Python + browser" trio has broad coverage, but on specific tasks it often loses to a toolset fine-tuned for the task (the paper's SWE-agent scoring higher on HumanEvalFix 1-shot is one piece of evidence). The trade between generality and point performance doesn't disappear because you built a platform.
  • The weight of the sandbox. One Docker container per session buys strong isolation and reproducibility at the cost of startup overhead, image size, and operational complexity (Docker-in-Docker, resource quotas). For local lightweight use, it's clearly over-engineered.
  • The complexity tax of being a "platform." Event stream, runtime, skill library, delegation, evaluation — each is an independent subsystem, and the bar for full understanding is high. Teams that just want to bolt an agent onto their product usually lose more than they gain by copying OpenHands' entire architecture — what's worth copying is its layered thinking.
  • Instability amid evolution. From OpenDevin to OpenHands to the V1 SDK / Agent Canvas, the project's APIs and directory structure change frequently; secondary development on top of it means tracking versions closely.

In one sentence: OpenHands proved that a harness can be systematically platformized and reproducibly studied — but it is itself neither the "lightest" nor the "strongest" agent. It's the instrument that makes "lightest" and "strongest" measurable.

Further reading ​

References ​