Skip to content

OpenHands

At a glance From OpenDevin to OpenHands: dissecting this open-source, community-driven full-stack development agent platform—the event stream architecture, the CodeAct "actions as code" action space, the Docker sandbox runtime, microagents, plus its real SWE-bench performance and its design divergence from SWE-agent.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

OpenHands ​

OpenHands (formerly OpenDevin) is currently the open-source world's closest thing to a "complete product" among software development agent platforms: it has a Web UI, a CLI, a hosted cloud version, and a programmable SDK, its core runs under the MIT license, and by mid-2026 its GitHub star count had passed 75K. It is maintained by the company All Hands AI but still governed as a community project—v3 of the platform paper (April 2025) lists more than 20 named authors spanning UIUC, CMU, and industry.

This page is not a usage tutorial (the docs site covers that thoroughly); instead it answers three questions engineers care about: why is its architecture shaped this way? What exactly is good about CodeAct's "actions as code" idea? And where does it philosophically diverge from an academically bred competitor, SWE-agent?

1. From OpenDevin to OpenHands: An Open-Source Community Strikes Back ​

The timeline is worth remembering because it is a microcosm of the 2024 agent boom:

  1. March 2024: Cognition released Devin, billed as "the first AI software engineer"—fully closed source and invite-only, and the community echoed with "I could build this too."
  2. March 2024 (within weeks): UIUC PhD student Xingyao Wang and others started the OpenDevin project as an open-source replica of Devin; its GitHub stars exploded in the first weeks, making it one of the fastest-growing open-source projects of the year.
  3. During 2024: the project renamed to OpenHands; the founding team incorporated All Hands AI, with CEO Robert Brennan (ex-Google, former Fairwinds open source lead), Graham Neubig (CMU professor) as chief scientist, and Xingyao Wang as co-founder.
  4. November 2025: the OpenHands Software Agent SDK released (paper arXiv:2511.03690), a thorough refactor of the agent components (internally called V1) that split "the platform" from "the embeddable SDK."
  5. 2026: the repo moved to a standalone OpenHands GitHub organization; the main project entered a 1.x release cycle (1.5.0 in March 2026), with the SDK versioning independently (at v1.31.x by July 2026).

Quick Facts (as of August 2026)

  • GitHub: 75K+ stars, thousands of forks; cumulative downloads exceeded 4 million by the end of 2025
  • Papers: the OpenHands platform paper (arXiv:2407.16741), the CodeAct paper (arXiv:2402.01030, ICML 2024), and the Agent SDK paper (arXiv:2511.03690)
  • Funding: $5M seed (September 2024, led by Menlo Ventures); $18.8M Series A (November 2025, led by Madrona)

Its positioning in one sentence: whatever Devin can do, it can mostly do too—and you can see every line of the implementation and swap out any part. That determines its popularity with two groups: teams that want to deploy agents on their own infrastructure, and engineers and researchers who want to study how a development agent is actually built.

2. Architecture: Event Stream, CodeAct, and the Sandbox Runtime ​

The OpenHands platform paper (arXiv:2407.16741) splits the architecture into three building blocks: the Agent abstraction, the Event Stream, and the Agent Runtime. Understand these three and you understand the whole system.

                 ┌────────────────────────────────────┐
                 │          Frontend / CLI / API      │
                 └──────────────┬─────────────────────┘
                                 │ user message
                 ┌──────────────▼─────────────────────┐
                 │         AgentController            │
                 │  state machine: RUNNING/AWAITING/  │
                 │  ERROR...                          │
                 └──────────────┬─────────────────────┘
                                 │ step()
┌───────────────────────────────┼───────────────────────────────┐
│                  Event Stream (append-only)                    │
│  Action ──► execution ──► Observation ──► appended to the     │
│  stream ──► next decision                                     │
│  (CmdRun / IPythonRunCell / FileEdit / Browse / Message ...)  │
└───────────────┬───────────────────────────────┬───────────────┘
                │ actions dispatched to         │
        ┌───────▼───────────────────────────────▼────────┐
        │           Agent Runtime (Docker sandbox)       │
        │  bash session │ Jupyter kernel │ browser │ fs  │
        └─────────────────────────────────────────────────┘

The Event Stream: Everything Is an Event ​

OpenHands has no "conversation history"—only an event stream. User messages, the agent's thoughts, every action (Action), every execution result (Observation) are all events appended to the same stream. At each step, the agent's input is a view of the event stream and its output is a new action.

The payoff of this design shows up in three places:

  • Replayable and debuggable. Once a task finishes, the whole trajectory is the event stream itself, exportable as-is and replayable item by item—also the foundation of its evaluation and observability work.
  • Externalized state. The agent itself is nearly stateless; pausing, resuming, and human handoff (a human inserting a message into the stream at any time) are all natural operations.
  • Multi-agent composes naturally. Delegation is simply the parent agent emitting an AgentDelegateAction; the child agent works under the same event-stream semantics and returns results as an Observation.

The V1 SDK went further, turning the event stream into a formal event sourcing implementation. The SDK paper's production-data conclusion: compared to V0, V1 significantly reduced the failure rate "attributable to the system itself," while event sourcing's overhead was negligible—an empirical validation that "engineering abstractions don't slow agents down."

CodeAct: A Unified Action Space ​

The default CodeActAgent does not define a pile of JSON function-calling tools; instead it gives the model a few "work like a human" capabilities: run bash commands, execute Python cells in Jupyter, read/write/edit files, operate a browser, message the human. The model invokes these capabilities by writing code. This is the platform's most central and most idea-dense design; Section 3 unpacks it.

The Runtime Sandbox ​

All actions execute inside an on-demand Docker container: a persistent bash session, a Jupyter IPython kernel, a headless browser, with the task workspace mounted. Two engineering corollaries follow:

  • A clear security boundary. The agent has full shell privileges, but the blast radius is confined to the container. The sandbox is not optional—letting LLM-generated code run directly on the host is handing the rm -rf trigger to a system that hallucinates; see Agent Security.
  • Environment as image. Evaluation (e.g., SWE-bench) and enterprise deployment reuse the same mechanism: package the task environment as a Docker image and let the agent work inside it. Environment reproducibility is a major source of credibility for OpenHands' evaluation results.

Microagents ​

Note that "microagent" has two historical meanings in OpenHands; don't confuse them:

  • In the paper: the platform paper's micro-agent refers to lightweight, single-purpose agent implementations—fixed prompt + fixed logic completing a narrow task (like submitting a PR or summarizing a discussion)—used to prove that "not every task needs the full CodeAct loop."
  • In the current product: Markdown files under the .openhands/microagents/ directory, essentially domain knowledge injected on demand—knowledge files with trigger keywords (injected only when a term is mentioned) and repo-level common-knowledge files (like AGENTS.md; see writing-agents-md). They belong to context engineering, not independently running agents.

A typical knowledge microagent looks like this—the frontmatter controls triggering, and the body is the injected context:

markdown
---
name: project-conventions
triggers:          # injected only when these keywords are hit; not always in context
  - tests
  - test
---

# Testing conventions for this repo
- Unit tests go in tests/unit/, using pytest; unittest is not allowed
- After changing code, `make test-fast` must pass before committing
- Fixtures live in tests/conftest.py; don't redefine them in new files

What's worth borrowing from this mechanism is its laziness: knowledge enters the context window only when relevant, avoiding stuffing the whole team wiki into the system prompt. When you build your own agent, "dynamically assembling context by keyword/task type" usually beats "one giant system prompt" on both token cost and accuracy.

3. CodeAct in Depth: Why "Actions as Code" ​

CodeAct comes from the ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" (arXiv:2402.01030, Xingyao Wang et al., UIUC + Apple). It is the theoretical source of OpenHands' action space and one of the most practically influential agent design ideas of recent years.

The Problem: JSON Tool Calling's Two Ceilings ​

Mainstream agent frameworks have the model emit predefined-format JSON or text to call tools (see Tools & MCP). The paper identifies two structural limits:

  1. Restricted action space: the model can only invoke tools you predefined; its compositional ability beyond the tool list is zero.
  2. No control flow or data flow: a JSON call is an atomic, stateless tool invocation. To "run a check on each of 10 files and aggregate the failures," you need 10 back-and-forth turns, with intermediate results hauled through the context in natural language.

The Solution: Let the Model Write Executable Python Directly ​

CodeAct unifies all of the agent's actions on the environment into a piece of executable Python code. Each turn: the model outputs code → the interpreter executes it → the result (including errors) comes back as an observation → the model corrects or emits a new action accordingly. The benefits stack:

  • Control flow and data flow: for-loop batch processing, if branches, storing intermediate results in variables for reuse—one block of code replaces a dozen-plus turns of the JSON approach. The paper's example: chaining multiple tools over a batch of inputs, feeding each tool's output to the next, done in a single CodeAct action.
  • Free access to the whole software ecosystem: the model can directly import pandas or call sklearn, without you hand-wrapping a tool per task.
  • Built-in debugging feedback: Python error messages are natural-language feedback designed for human programmers; the model using them for self-debugging comes naturally.
  • Aligned with the pretraining distribution: modern LLMs have seen far more code in pretraining than any proprietary JSON calling format; writing code is the model's "native tongue," at nearly zero adaptation cost.

The Evidence: Not Just Plausible-Sounding ​

The paper ran controlled experiments across 17 open and closed models:

ExperimentComparisonResult
API-Bank atomic callsCodeAct vs JSON vs text formatCodeAct on par or better on most models; the advantage is especially clear on open models (JSON is a weak spot for open models)
M3ToolEval (82 complex tasks requiring multi-tool, multi-turn composition)SameCodeAct's success rate up to 20 percentage points higher; interaction turns reduced by up to ~30%
Strongest model on the same benchmark (GPT-4-1106)CodeAct 74.4% vs text 53.7% vs JSON 52.4%A gap of more than 20 percentage points

The paper also did something with deeper reach: it built CodeActInstruct, an instruction fine-tuning dataset of 7K CodeAct multi-turn interaction trajectories, and trained CodeActAgent (on Llama-2 7B / Mistral 7B), proving that "code as action" can be trained into a model, not just prompted out of one. This directly inspired a wave of later agentic fine-tuning work.

A Judgment Worth Remembering

CodeAct's insight had become industry consensus by 2026: Claude Code's Bash tool and the various "model writes shell directly" coding agents are all CodeAct thinking at heart. But mind the boundary—it holds only when the model is strong enough and there is a sandbox. Weak models produce far higher error rates generating free code than with constrained JSON; without an isolated environment, free code execution is an unacceptable risk. JSON function calling remains the right answer for controlled enterprise integration (permission auditing, schema validation).

4. Evaluation: OpenHands on SWE-bench ​

OpenHands is a fixture of the SWE-bench ecosystem; when reading its scores, watch three things: which benchmark version, paired with which model, and whether it's the official reproduction.

  • SWE-bench Verified (the 500-instance human-curated version) is the main battleground. OpenHands officially reported ~53% for the CodeAct solution in early 2025—a number that stirred controversy in community reproduction (GitHub issue #10767 specifically discusses the default config not reproducing it), a textbook case of "discount vendor-reported benchmark scores." Methodology: Evaluation.
  • On the official swebench.com leaderboard, OpenHands' subsequent entries kept improving: the August 2025 submission reached 69.6%–71.8%. Note that leaderboard scores are joint "scaffold + model" results; the top spots change hands with every new model release, so watching the trend beats memorizing numbers.
  • The platform paper's cross-comparison is more informative: it evaluated 15 benchmarks on the same runtime (SWE-bench, WebArena, GPQA, HumanEvalFix, etc.), and CodeActAgent was the strongest among open-source entries on the vast majority of coding and browsing tasks—proving the platform's generality rather than single-benchmark gaming.

Leaderboard-Reading Discipline

SWE-bench Verified went through a sprint from ~50% to 80%+ in 2025–2026; the top of the board is occupied by frontier models + heavily tuned scaffolds, and score gaps have entered statistical-noise territory. When selecting OpenHands, don't cite its leaderboard rank from a year ago—scaffolds iterate extremely fast, and the same scaffold with a different model can swing 20 points. The right posture: small-scale field tests with your own repos and tasks.

5. All Hands AI: The Company and Commercialization ​

OpenHands takes a textbook open-core path, executed with restraint:

  • Fully open core, MIT license. The platform, SDK, and runtime all live in the open repo; self-hosting requires no license fees—only your own LLM API bill.
  • Monetization via hosted and enterprise editions. OpenHands Cloud offers a no-ops hosted version: a free individual tier to start ($20 credit at signup; bring your own API key or use platform-supplied models at cost, no markup), with Slack, Jira, and GitHub/GitLab/Bitbucket integrations; the Team and Enterprise tiers sell organizational controls, SSO, private deployment, and fleet management—pricing via sales.
  • The company feeds the project; the project funnels to the company. The $18.8M Series A (November 2025, led by Madrona, with Menlo Ventures and others participating) mostly funds engineering of the Cloud product and SDK. The seed round (September 2024, $5M, led by Menlo) had a very open-source-flavored investor list: Hugging Face co-founder Thom Wolf, PyTorch author Soumith Chintala, and Cloudera co-founder Jeff Hammerbacher all in it.

For users, this structure is fairly healthy today: an MIT core plus a clear company funding channel is more trustworthy than "open source to gather a community, then change the license." But one detail is worth noting—the enterprise directory in the repo is separately licensed; check the LICENSE layout before commercial secondary development.

6. Ecosystem and Value for Secondary Development ​

If you want to "build your own agent standing on giants' shoulders," OpenHands is one of the best bases in 2026, for four reasons:

1. The SDK Split the Platform into Libraries ​

The Software Agent SDK released in November 2025 (OpenHands/software-agent-sdk) was the turning point. Before V1, reusing OpenHands' capabilities essentially meant forking the whole platform; the SDK extracts agent, tool, runtime, and event stream into standalone Python packages, with the paper stressing that a minimal agent implementation takes "just a few lines of code by default," while retaining production features like custom tools, memory management, and security analysis. Compared to the OpenAI Agents SDK, Claude Agent SDK, and Google ADK, its differentiation is natively integrated sandboxed execution, model-agnostic multi-LLM routing (100+ models via LiteLLM), and seamless local-to-remote execution migration. A minimal agent with it looks like:

python
# Minimal example based on the OpenHands Agent SDK
# (API per the package docs; it iterates quickly between versions)
from openhands.sdk import Agent, LLM, Conversation
from openhands.tools import BashTool, FileEditorTool

llm = LLM(model="anthropic/claude-sonnet-4-5")  # routed via LiteLLM; changing models is a one-line edit
agent = Agent(llm=llm, tools=[BashTool(), FileEditorTool()])

# Conversation wraps the event stream: send messages, run the loop, read event history
convo = Conversation(agent=agent, workspace="./my-project")
convo.send_message("Fix all failing unit tests under src/")
convo.run()

2. The Runtime Works Standalone ​

Even if you don't want its agent, the Docker sandbox runtime is independently valuable: many teams use it as "a safe execution environment for their own agents," saving the considerable grunt work of building a sandbox from scratch.

Deployment shapes at a glance: OpenHands in 2026 offers a fuller set of access surfaces than most peers; decide which one you need before choosing it as your base:

ShapeFitsCost
Local GUI (self-hosted Docker)Personal trial, small-team intranet deploymentYou manage images, model keys, and upgrades yourself
CLI / headless modeCI integration, batch jobs (e.g., overnight issue fixing)No UI; weak interactivity
Python SDK (software-agent-sdk)Embedding agent capability into your own productFast iteration; breaking API changes
OpenHands CloudTeams that don't want to run ops and need Slack/Jira integrationData leaves your domain; enterprise compliance requires the Enterprise tier

3. Evaluation Infrastructure Out of the Box ​

The evaluation harnesses for SWE-bench, WebArena, and other benchmarks are built into the repo, with environments imaged and trajectories exportable. For agent research or evals in practice, this is ready-made scaffolding.

4. The Community Activity Is Real ​

Weekly release cadence, hundreds of contributors, continuous papers—platform paper v3 records over 188 contributors and 2,100+ commits. For secondary developers, this means "the problem you hit has probably been hit before." On the flip side, be prepared: fast iteration means fast API drift; breaking changes on upgrade are the norm, and pinning versions is basic hygiene.

When to Pick OpenHands as Your Base

A fit when: you need self-hosting, arbitrary model swapping, changes to agent internals, or embedding the agent into your own product. Not a fit when: you just want to quietly code in a terminal (Claude Code or Cursor is less hassle), or you only need a single-file issue-fixing script (SWE-agent/mini-SWE-agent is lighter).

7. Design Comparison with SWE-agent ​

OpenHands and SWE-agent are the two poles of open-source SWE agents, and their papers are also this field's top two by citations. Their difference is not "which is stronger" but a divergence of two design philosophies:

DimensionOpenHandsSWE-agent
Origin and goalA product-grade platform replicating Devin, aiming at "a complete, usable AI developer"A Princeton academic research project, aiming to answer "which interfaces let an LM solve real issues"
Core abstractionEvent stream + CodeAct: actions as code, running bash/Python directlyACI (Agent-Computer Interface): custom commands designed for the LM (open/edit/lint, etc.)—interface design for the model the way UI is designed for humans
Trajectory shapeEvent stream, supporting delegation, multi-agent, and human intervention anytimeLinear "thought-action-observation" loop, deliberately kept simple; mini-SWE-agent cuts it to ~100 lines with bash only
InterfacesWeb GUI, CLI, VS Code, Cloud, SDKMostly command line, research-oriented
Engineering weightHeavy: Docker runtime, frontend, server, SDK; deployment has a barLight: a single process runs a whole trajectory; easy to reproduce and modify
Typical usersEngineering teams deploying/embedding/customizing the productPeople doing agent research and large-scale evaluation experiments

The most interesting divergence is the action space. The SWE-agent ACI paper argues: giving the model a set of carefully designed custom commands is more reliable than giving it a raw shell—because the model's shell knowledge is noisy, and custom commands narrow the error surface. The CodeAct argument is exactly the opposite: give the model free code execution, fully releasing control flow and ecosystem libraries, with reliability underwritten by the sandbox and error feedback.

Who's right? Looking back from 2026, the answer is "both were partly right, and the market is converging":

  • Under the premise of strong models + sandboxes, the CodeAct route clearly wins; industry products (Claude Code, Codex, etc.) have basically all converged on "the model uses bash directly";
  • But the ACI insight—interface design sets the ceiling of agent performance—was fully inherited: OpenHands' file editing tool likewise went through rounds of ACI-style polish (from free-form sed to editor commands with line-number validation), and mini-SWE-agent solving SWE-bench with pure bare bash also proved the floor of "simple interface + strong model" can be quite high.

The real conclusion: there is no silver bullet in action-space design; it is a function of "model capability × task type × safety constraints." OpenHands chose freedom first, backstopped by engineering; SWE-agent chose constraint first, backstopped by simplicity. To experience both philosophies hands-on, the cheapest way is to run a few SWE-bench instances on each and read the trajectories line by line—more eye-opening than ten comparison articles; see Build Your Own Agent.

References ​