Appearance
Case Study: OpenHands
If SWE-agent demonstrated how powerful a carefully polished ACI (agent-computer interface) can be, OpenHands answers a different question: can the harness itself be built as a general-purpose, reproducible experimental platform — where anyone can swap agents, models, and tools, and run benchmarks, all on the same infrastructure?
OpenHands (formerly OpenDevin) is probably the most thorough embodiment to date of "the harness as a research object." It isn't an implementation of some agent — it's a complete scaffold around agents: an event stream, a sandboxed runtime, a skill library, multi-agent delegation, an evaluation framework. Understanding it is roughly equivalent to understanding the full checklist of a modern coding agent harness.
The project at a glance
| Dimension | Details |
|---|---|
| Former name | OpenDevin, started March 2024 as an open-source replica of Devin; renamed OpenHands in September 2024 |
| Paper | OpenHands: An Open Platform for AI Software Developers as Generalist Agents (arXiv:2407.16741, v1 2024-07-23) |
| Core authors | Xingyao Wang, Graham Neubig, et al. (UIUC / CMU / All Hands AI, among others) |
| Theoretical basis | The CodeAct paper: Executable Code Actions Elicit Better LLM Agents (arXiv:2402.01030, ICML 2024) |
| License | MIT (commercial use allowed) |
| Community scale | At publication: 32K GitHub stars, 188 contributors, 2.1K+ contributions; still growing since |
| Positioning | A general agent development platform: agent abstraction + event stream + sandbox runtime + AgentHub + eval framework |
A note on project evolution
OpenHands iterates fast. This page dissects its classic architecture (matching the 2407.16741 paper and the 0.x-era code structure). In November 2025 the project released V1 along with a standalone OpenHands Software Agent SDK; the repo moved from All-Hands-AI/OpenHands to OpenHands/OpenHands, and the product expanded into Agent Canvas, which can host third-party agents (Claude Code, Codex, etc.). The architectural ideas carry through, but for specific class names and directories, defer to the version of the code you're reading.
Why OpenHands is worth studying
Most coding agent products (Claude Code, Cursor, Aider) hide the harness inside the product: you can observe behavior and outcomes, but not the full mechanism. OpenHands does the opposite:
- All mechanisms in the open. Event stream, state, actions, observations, runtime — every concept has a clear code entity, and is written up in the paper.
- Controllable variables. The agent implementation, the LLM, the sandbox image, the skill library, the benchmarks — all are swappable slots, naturally suited to ablation experiments.
- Academic backing. Two papers (the platform paper + the CodeAct paper) spell out the motivation behind design decisions — no reverse engineering required.
In the language of the anatomy page: OpenHands is the project that made every harness component a "first-class citizen." Its own positioning is blunt — a community-driven platform, not an agent.
Overall architecture: the big three
The platform paper condenses OpenHands into three main components: the agent abstraction (community-contributed agent implementations, gathered in AgentHub), the event stream (a history of actions and observations), and the runtime (executes actions into observations).
┌─────────────────────── OpenHands host process ───────────────────────┐
│ │
│ User / UI / CLI │
│ │ MessageAction / feedback │
│ ▼ │
│ ┌─────────────┐ step(state) ┌──────────────────────┐ │
│ │ Agent │ ◄───────────────── │ State │ │
│ │ (CodeAct et │ │ ├ event stream │ │
│ │ al.) │ │ │ (action+obs │ │
│ └──────┬──────┘ │ │ history) │ │
│ │ returns an Action │ ├ cumulative LLM cost │ │
│ ▼ │ └ delegation metadata │ │
│ ┌──────────────── Event Stream ────────────────┐ └──────────┘ │
│ │ Action → dispatched to Runtime → │ │
│ │ Observation written back to the stream │ │
│ └──────────────────────┬───────────────────────┘ │
└──────────────────────────┼────────────────────────────────────────────┘
│ REST API (action execution requests)
▼
┌─── Docker sandbox container (one per task session) ───┐
│ Action Execution API Server │
│ ├─ bash shell (command execution) │
│ ├─ Jupyter IPython server (runs Python) │
│ └─ Chromium + Playwright (browser) │
│ Mount: the user's workspace directory │
└───────────────────────────────────────────────────────┘The most striking thing about this structure is the separation of responsibilities: the agent only "looks at state and emits actions," with no idea how actions get executed; the runtime only "executes actions and returns observations," with no idea who emitted them or why. The two are decoupled by the event stream, an append-only log. Now, block by block.
The event stream: everything is an event
The event stream is OpenHands' heart. Its definition (paper §2.1): a time-ordered collection of actions and observations, including both the actions the agent emits and the user's instructions and feedback.
Several key designs revolve around the stream:
State = event stream + auxiliary data. Beyond the event history, state carries the cumulative LLM cost, tracking metadata for multi-agent delegation, execution parameters, and so on. This means harness-level concerns (how much money has been spent, who delegated which subtask to whom) and the agent's "memory" are explicitly modeled in the same data structure.
Actions and observations are paired typed objects. The core action set is deliberately small and general:
| Action | Purpose |
|---|---|
CmdRunAction | Execute an arbitrary bash command in the sandbox |
IPythonRunCellAction | Execute Python code in the sandbox's Jupyter |
BrowserInteractiveAction | Drive the browser via BrowserGym's domain-specific language |
AgentDelegateAction | Delegate a subtask to another agent |
AgentFinishAction / MessageAction | End the task / converse with a human |
The corresponding observations describe environment changes: command output, Python execution results, the browser's HTML/DOM/accessibility tree/screenshots, user messages, and so on.
The event stream is both the conversation history and the audit log. Because every interaction lands on the same stream, "what context the model saw" and "what actually happened in the system" are two views of the same thing. That makes debugging, replay, and evaluation direct — the eval framework consumes exactly these trajectories to score runs.
The harness view
The event-stream architecture is essentially event sourcing applied to agent systems: don't store "current state," store "the sequence of events," and derive state by replay. The cost is serializing the history into the prompt at every step — hence the pressure on context engineering (long-session compression and truncation are handled by a dedicated condenser mechanism). See Context Engineering.
The agent abstraction: one step function
On top of the event stream, the interface for implementing a new agent is compressed to the extreme — the MinimalAgent skeleton in the paper's Figure 3 says it all:
python
class MinimalAgent:
def reset(self) -> None:
self.system_message = "You are a helpful assistant …"
def step(self, state: State):
# 1. Serialize the event-stream history into a message list
messages = [{"role": "system", "content": self.system_message}]
for prev_action, obs in state.history:
messages.append(get_action_message(prev_action))
messages.append(get_observation_message(obs))
# 2. Call the LLM and parse the output into an action
response = self.llm.do_completion(messages)
action = self.parse_response(response)
# 3. Return a typed action for the runtime to execute
if self.is_finish_command(action):
return AgentFinishAction()
elif self.is_bash_command(action):
return CmdRunAction(command=action.command)
elif self.is_python_code(action):
return IPythonRunCellAction(code=action.code)
elif self.is_browser_action(action):
return BrowseInteractiveAction(code=action.code)
else:
return MessageAction(content=action.message)That's what the agent loop looks like in OpenHands: step(state) -> action, with the loop itself driven by the platform. The value of the abstraction: researchers only need to care about "given this history, what's the next step" — execution details, sandbox safety, cost accounting, UI rendering are all covered by the harness. That's the precondition for being an experimental platform: swapping agents costs about as much as writing a class a few dozen lines long.
CodeAct: code as the unified action space
OpenHands' default general agent, CodeActAgent, is based on the same team's earlier CodeAct paper (arXiv:2402.01030, ICML 2024). That paper's question: in what format should an agent's actions be expressed?
Both mainstream approaches of the time had obvious flaws:
| Action format | Problem |
|---|---|
| Predefined JSON function calling | The action space is capped by the pre-declared tool catalog; one action calls one tool, with no composition |
| Free-text actions | Fragile parsing, format drift, unreliable execution |
CodeAct's answer: have the LLM emit executable Python code as the action, run it through an interpreter, observe the result, continue. Code natively carries loops, conditionals, variables, library imports, and function composition — a single code block can accomplish what JSON actions would need a dozen round trips for. Better yet, the interpreter's error messages are themselves deterministic, high-quality environment feedback the agent can use to self-correct.
The paper tested 17 LLMs on API-Bank and a newly built tool-use benchmark: CodeAct improved success rates by up to about 20% over JSON/text formats.
OpenHands engineering-izes this idea:
- The core actions
IPythonRunCellActionandCmdRunActionare CodeAct instantiated — the action space ≈ "anything runnable on a Linux machine." - It also stays compatible with traditional function calling: to add a new "tool," you don't touch an action schema — you write a Python function and drop it into the sandbox's IPython environment (see AgentSkills, next section).
- The more radical corollary: the agent can build its own tools — when no ready-made API exists, it writes a Python function on the spot.
The price of a unified action space
"Everything is code" maximizes flexibility — and maximizes the safety problem: every line of model-written code gets really executed. That's exactly why OpenHands must pair with a strongly isolated sandbox runtime. CodeAct and the Docker sandbox are mutually presupposing designs on this platform, not two independent features.
The runtime sandbox: locking danger in a container
CodeAct lets the agent execute arbitrary code; the runtime's job is to make that safe and reproducible. Key points of OpenHands' runtime design (paper §2.2):
One container per session. Every task session starts an isolated Docker container, and every action in the event stream executes inside it. The user's working directory is mounted in with configurable options, capping the agent's blast radius at the container.
An in-container action-execution API. A REST API server runs inside the sandbox (the OpenHands Action Execution API); the host process sends it actions from the event stream, and it maintains three things and returns observations:
- A bash shell attached to the container's OS;
- A Jupyter IPython server handling interactive Python execution;
- A Playwright-based Chromium browser, whose action primitives come from BrowserGym (navigate, click, type, scroll), with observations including HTML, DOM, accessibility tree, screenshots, open tabs, and so on.
Arbitrary image support. The runtime has a build mechanism: take any user-provided Docker image, inject the action-execution API, and it becomes a usable sandbox. That means evaluating different projects (different languages, dependencies, OS environments) requires no harness changes — just swap images. Essential for benchmarks like SWE-bench that require reproducing an environment per repository.
Why does the browser go in the sandbox too?
OpenHands' ambition isn't "a code completion tool" but "a general digital worker": human developers consult docs, look at issues, and browse the web while working. The built-in browser lets the same agent abstraction cover web tasks like WebArena — in the paper, CodeActAgent runs software engineering, web browsing, and misc-assistant benchmarks simultaneously without any prompt changes, purely thanks to this action-space design.
AgentSkills: tools defined not by schema but by pip
The SWE-agent paper proved that carefully designed tools (the ACI) massively affect performance — but creating, maintaining, and distributing tools is heavy engineering. OpenHands' AgentSkills library offers a very light answer (paper §2.3):
- A skill is just a Python package. Writing a Python function is writing a tool; it's auto-imported into the sandbox's IPython environment and the agent calls it directly via
IPythonRunCellAction. No schema registration, no agent code changes. - Restrained inclusion criteria. The official rules admit only two kinds: things the LLM can't reliably do by writing code itself (like
edit_file, which edits files by line number), and things that require calling an external model (likeparse_imagecalling a vision model orparse_pdffor PDFs). "The LLM already knows how to read a CSV with pandas" doesn't need wrapping. - Maintained like real software. The skill library ships with unit tests, avoiding "the agent takes the blame, the tool stabs it in the back."
This route continues the "tools as code" discussion from the tool component page: when the action space itself is code, "adding a tool to the agent" degenerates into "installing a dependency in the environment," and the engineering complexity drops dramatically.
microagent: two confusable concepts
In OpenHands-speak, "micro" appears at two levels, easily conflated in the docs; worth separating:
1. Micro agent (the specialized agent in the paper). Reuses most of a general agent's implementation (like CodeActAgent), swapping only the prompt and a bit of configuration to specialize for a task. The design goal is lowering the contribution bar — you don't need to understand the whole harness; sharing a well-tuned prompt publishes a "new agent."
2. Microagents (the knowledge-fragment mechanism in the repo). This is the mechanism users actually touch: create a .openhands/microagents/ directory at the repo root, drop in Markdown files containing project-specific guidance for the agent. The official docs define two types:
| Type | Trigger | Frontmatter | Typical use |
|---|---|---|---|
| General | Injected automatically every session | Optional | repo.md: repo structure, build commands, code conventions |
| Keyword-Triggered | Injected only when a specified keyword appears in the user's message | Required (triggers field) | Specialized knowledge loaded only when a particular framework/service comes up |
some-repository/
└── .openhands/
└── microagents/
├── repo.md # general repo guidance (always on)
├── trigger_this.md # keyword-triggered
└── trigger_that.md # keyword-triggeredThe harness view
Microagents are essentially conditional context injection: turning "what knowledge the model should see, when" from a one-shot giant system prompt into a modular, on-demand knowledge base. The docs explicitly warn that loaded microagents occupy the context window — so keyword triggering isn't a nice-to-have, it's a necessary means of controlling prompt size. Same idea, different implementation from Claude Code's skills and layered CLAUDE.md loading.
Multi-agent delegation and the pluggable design
Delegation is an action, not a framework. Multi-agent collaboration in OpenHands works through AgentDelegateAction: one agent packages a subtask and delegates it to another — for example, the general CodeActAgent delegating complex web operations to the dedicated BrowsingAgent. Delegation relationships are recorded in State's metadata. The design is deliberately simple — no blackboard system, no role-orchestration engine, just "one more action in the action space." See the comparative discussion in multi-agent orchestration.
Pluggable agents: AgentHub. The platform gathers 10+ community-contributed agent implementations: the default CodeActAgent, the web-focused BrowsingAgent, GPTSwarm's optimizable-graph agent, various micro agents… They all share the same event stream and runtime, so comparisons are fair.
Pluggable LLMs: LiteLLM. Model access goes through LiteLLM as a unified gateway; switching models means changing the model string in config (e.g. openai/gpt-4o, claude-*, OpenRouter, or local vLLM/Ollama endpoints), with practical logic like 429 rate-limit retries built in. Model and harness are fully decoupled — a living specimen of this site's model vs. harness argument.
Pluggable evaluation. The platform integrates 15 benchmarks (SWE-bench, WebArena, GAIA, GPQA, HumanEvalFix, ML-Bench, BIRD, AgentBench, MINT, etc.) across the three categories of software engineering, web browsing, and misc assistant. Representative numbers from the paper: CodeActAgent v1.8 with claude-3.5-sonnet reached 26.0% on SWE-bench Lite (average cost per instance about $1.10); the same agent, without prompt changes, reached 14.5–15.3% on WebArena and about 52% on the GPQA diamond subset.
A research finding worth noting
On AgentBench, the paper observed that when switched to a weaker model (gpt-3.5-turbo), OpenHands' general agent actually did worse than the specialized baselines in the original benchmark. The authors' interpretation: general agents have a threshold requirement on the base model's instruction-following ability; below that threshold, a carefully specialized dedicated harness is more stable. Quantitative evidence for "matching the harness to the model's capability."
Trade-offs and critical observations
OpenHands is an excellent research platform, but every design choice has a matching cost; see them clearly before making engineering decisions:
- The full-serialization cost of the event stream. Sending the complete history to the model at every step inflates token consumption and latency on long tasks, forcing reliance on the condenser for compression — and the compression strategy itself becomes a new research variable. For production systems, append-only event sourcing isn't necessarily the best answer to context management.
- The tension between a general action space and a bespoke ACI. The "bash + Python + browser" trio has broad coverage, but on specific tasks it often loses to a toolset fine-tuned for the task (the paper's SWE-agent scoring higher on HumanEvalFix 1-shot is one piece of evidence). The trade between generality and point performance doesn't disappear because you built a platform.
- The weight of the sandbox. One Docker container per session buys strong isolation and reproducibility at the cost of startup overhead, image size, and operational complexity (Docker-in-Docker, resource quotas). For local lightweight use, it's clearly over-engineered.
- The complexity tax of being a "platform." Event stream, runtime, skill library, delegation, evaluation — each is an independent subsystem, and the bar for full understanding is high. Teams that just want to bolt an agent onto their product usually lose more than they gain by copying OpenHands' entire architecture — what's worth copying is its layered thinking.
- Instability amid evolution. From OpenDevin to OpenHands to the V1 SDK / Agent Canvas, the project's APIs and directory structure change frequently; secondary development on top of it means tracking versions closely.
In one sentence: OpenHands proved that a harness can be systematically platformized and reproducibly studied — but it is itself neither the "lightest" nor the "strongest" agent. It's the instrument that makes "lightest" and "strongest" measurable.
Further reading
- Case Study: SWE-agent — the contrasting sample for ACI thinking: a small, sharp, dedicated harness
- Case Study: Devin — the commercial product OpenDevin originally benchmarked itself against
- Case Study: Aider — another open-source route, one that polishes the "edit tool" to the extreme
- The Agent Loop — the general form of
step(state) -> action - Context Engineering — how the event stream becomes the prompt, and the compression problem
- Tools & MCP — CodeAct vs. function calling: the battle of action spaces
- Permissions, Safety & Human-in-the-Loop — beyond the sandbox: designing "when to stop and ask a human"
- Classic Papers, Annotated — more background on the CodeAct and platform papers
- Build a Minimal Harness from Scratch — putting this page's layered thinking into a few hundred lines of code
References
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents (arXiv:2407.16741)
- Executable Code Actions Elicit Better LLM Agents (arXiv:2402.01030, ICML 2024)
- The OpenHands GitHub repository
- OpenHands docs: Microagents overview
- OpenHands docs: LLM configuration (LiteLLM)
- OpenHands docs: Architecture overview