Skip to content

Devin

At a glance A full retrospective of Cognition's "AI software engineer" Devin: from the sensational 2024 launch at 13.86% on SWE-bench and the Upwork demo controversy, through the 25x price cut of 2025 and the Windsurf acquisition, to the 2026 $26B valuation and $492M annualized revenue—plus the real lessons it left for agent engineering.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Devin ​

On March 12, 2024, a small company called Cognition released a demo video announcing that it had built "the world's first AI software engineer," Devin. The founding team included International Olympiad in Informatics (IOI) gold medalists; behind them stood a $21 million Series A led by Founders Fund and a string of Silicon Valley luminaries including the Collison brothers and Elad Gil. In the video, Devin reads documentation on its own, sets up its own environment, debugs itself, and even "takes real freelance jobs on Upwork and earns money."

What happened over the following two and a half years nearly condenses the growth story of the entire agent industry: overnight deification, frame-by-frame debunking, quiet polish, quiet commercialization, a blockbuster acquisition, a valuation rocket. As of May 2026, Cognition had raised over $1 billion at a $26 billion valuation, annualized (run-rate) revenue reached $492 million, and the customer list included Citigroup, Mercedes-Benz, Goldman Sachs, and the U.S. Army and Navy.

This page neither repeats the myth nor recycles the mockery; it dissects Devin as a specimen: what it did, which parts were real, and which lessons everyone building agents should remember.

1. The Launch: Five Weeks from Sensation to Debunking ​

1.1 The March 2024 Shock ​

Devin's hardest evidence at launch was its SWE-bench score: 13.86% of real GitHub issues solved end to end, versus a best unassisted baseline of 1.96%—even when the model was told which files to modify, the best score at the time was only 4.80%. A roughly 7x gap, and unassisted versus assisted.

But note the footnotes in the official technical report—this 13.86% came with three qualifiers:

  • It was evaluated only on a random 25% subset of the full SWE-bench set;
  • Most other models were compared under "told which files to modify" (assisted) conditions, while Devin was unassisted—this favors Devin and is a genuine plus;
  • The evaluation was run by Cognition itself, with no third-party replication.

To be fair, 13.86% in March 2024 was genuine SOTA, and it directly ignited the arms race around "agents solving issues end to end." SWE-bench subsequently became the de facto industry standard; the methodological discussion around it lives in Agent Evaluation and the SWE-agent case study.

The intensity of the discourse still looks astonishing in hindsight: headlines like "software engineers gone in five years" and "CS degrees are worthless" filled tech media; Devin's waitlist codes were hard to come by, and secondhand invites became status symbols on social platforms. That fervor was itself the fuel for the later backlash—pushing expectations to an unfulfillable height was Devin's PR strategy's greatest success and its greatest mistake.

1.2 Five Weeks Later: The Frame-by-Frame Debunking ​

In April 2024, Carl Brown, a developer with 35 years of engineering experience, published a 25-minute video on his YouTube channel Internet of Bugs: "Debunking Devin: 'First AI Software Engineer' Upwork lie exposed!" It flooded Hacker News and Reddit. He took apart the segment of Cognition's official demo where "Devin completes a real freelance job on Upwork," frame by frame, and found several hard problems:

  • The task was carefully chosen and mischaracterized. The Upwork client wanted help getting an existing road-damage-detection vision model running—an environment setup and ops problem at heart—yet the demo presented Devin as "writing code from scratch to complete a development task."
  • Devin was fixing bugs it had itself created. Many of the errors Devin spent so long debugging in the video were introduced by its own generated code—create problems, then solve them, and it looks like work.
  • No real Upwork process took place. No bidding, no delivery, no evidence of client acceptance or payment.
  • The output code was poor. Unnecessarily complex and redundant; a human reproduced the same result in about 36 minutes, while Devin flailed for over 6 hours.

The industry significance of this episode far exceeds its gossip value: it established the concrete archetype of the term "demo video trap"—an agent demo never shows the distribution of capability; it shows the best sample from the distribution of luck. The problem persists: in 2026, vendor launch events still run the same play.

The Demo Video Trap

When watching any agent product demo, ask three questions by default: how many runs did it take to succeed once? Who picked the task? What is the human cleanup cost after failure? The lesson of the Devin affair is not "this company is a fraud" but "a single demonstration carries roughly zero evidentiary weight." Real signal comes from third-party, long-horizon field tests—like Answer.AI's below.

1.3 A Year Later: Third-Party Field Test—20 Tasks, 14 Failures ​

In January 2025, the Answer.AI team (the lab founded by Jeremy Howard) published "Thoughts On A Month With Devin"—a rare, detailed review based on a month of real use. They assigned Devin 20 real tasks across four categories: new projects, research, analyzing existing code, and modifying existing code. The results:

OutcomeCountNotes
Succeeded3Including "glue tasks" like the initial Notion→Google Sheets data migration
Failed14Overcomplicated "code soup," stuck in dead loops, hallucinated nonexistent features
Inconclusive3Produced output but didn't truly solve the problem

More damning than the failure rate was the unpredictability: "We could not find any pattern to predict which tasks would succeed." Plus the classic symptom of autonomy turning into liability—asked to deploy multiple applications to a single Railway instance (something Railway doesn't support at all), Devin failed to recognize the task was infeasible and instead spent a whole day trying approaches and hallucinating nonexistent features.

This review and the April 2024 debunking video together form the discursive backdrop of Devin's first act. Worth stressing: they criticized the version of Devin from late 2024 to early 2025; as we'll see, both product and market changed materially afterward.

2. Product Shape: A "Remote Colleague" Living in the Cloud ​

The fundamental difference between Devin and Cursor or Claude Code is not model capability but interaction topology. It is not a copilot in your IDE, nor a command-line companion in your terminal—it is a remote employee with its own desk:

      You (human)                       Devin (cloud)
┌────────────────────┐      ┌──────────────────────────────────┐
│  Slack / Web UI    │      │  Dedicated cloud sandbox         │
│  ─ Assign tasks    │      │  (Docker container)              │
│  ─ Watch progress  │─────►│  ├─ shell (install deps, run     │
│  ─ Correct course  │      │  │   tests)                      │
│  ─ Review PRs      │      │  ├─ Code editor                  │
└────────────────────┘      │  ├─ Browser (read docs, preview  │
                            │  │   web pages)                  │
                            │  └─ Long-task planner            │
                            │           │                      │
                            │           ▼                      │
                            │   Submits a Pull Request when    │
                            │   done                           │
                            └──────────────────────────────────┘

This shape brings three structural traits:

  • Asynchronous long tasks. You assign a task and go to a meeting; it runs for hours or even days. This is the essential difference from Cursor (synchronous, human watching every line) and Claude Code (synchronous, authorizing every change).
  • Ships with a complete environment. The sandbox has shell, editor, and browser; it can install dependencies, read API docs, and preview the web pages it writes. The environment's completeness is its moat—and also the amplifier of its failures: in Answer.AI's field test, Devin often burned its budget on environment problems.
  • Delivers via PR. The output of a job is a Pull Request entering the existing code review process. This is crucial: Devin didn't ask engineers to change how they accept work—it embedded itself into the existing engineering collaboration structure.

2.1 Product Line Expansion ​

After Devin 2.0 shipped in April 2025, the product boundary visibly expanded from "a single remote engineer" to "AI infrastructure for an engineering team":

  • Parallel Devins: dispatch multiple Devin instances simultaneously to work through a task queue in parallel, with an overview UI to manage them. Management granularity shifted from "watching one agent" to "running a fleet."
  • Devin Wiki: automatically generates and maintains documentation for your repo—architecture diagrams, source links, module descriptions, continuously updated as the code evolves. It aims at the most painful enterprise problem: existing knowledge.
  • Devin Search: semantic Q&A over the codebase, with cited answers.
  • Interactive Planning: aligns on a plan with you before starting work, splitting fuzzy requirements into confirmable steps before executing. A direct response to the "autonomy liability" problem—moving the correction point earlier.
  • Agent-native IDE: a VS Code-like interface where you can jump in, take over, or co-edit at any point while Devin works.
  • API & integrations: Slack, a VS Code extension, and an API, letting it be embedded into CI/CD and ticketing systems for automatic triggering.

Note the implicit logic of this evolution: Devin in 2024 sold "replacing engineers"; Devin from 2025 on sells "engineering leverage"—Wiki and Search aren't even about writing code, they sell codebase understanding. This pivot is the key to its survival, unpacked below.

2.2 The Lifecycle of a Typical Task ​

Turning the abstract architecture into concrete operations, a Devin task looks roughly like this:

  1. Assign the work: describe the task in Slack with @Devin (or in the web UI), attaching issue links, relevant docs, and environment hints. Task description quality directly decides success or failure—this is the part of using Devin that most resembles onboarding a new hire.
  2. Align on the plan (post-2.0): Devin first outputs an execution plan, breaking out steps and listing assumptions, and waits for your confirmation or corrections before starting. Skipping this step and letting it run is the most common rookie failure mode.
  3. Execute asynchronously: it clones the repo, installs dependencies, runs tests, and writes code in its own sandbox, reporting progress in Slack as it goes. You can interject corrections at any time, or open the IDE interface and take over directly.
  4. Deliver and accept: when done it opens a PR with a summary of its work. Then it enters your team's normal code review—no special channel; it is treated as an ordinary contributor.
  5. Billing: charged by the ACUs actually consumed. A failed task stuck for a full day and a successful task done in ten minutes can cost about the same—which makes "killing runaway tasks early" the core survival skill of using Devin.

Point 5 is many teams' first lesson: cost control with Devin is fundamentally the management discipline of task granularity and stop-losses, not a technical problem. See Cost Engineering for methodology.

3. Public Technical Information: Sparse, but Substantive ​

Cognition keeps architecture details tightly guarded, but a few pieces of public information are worth close reading.

3.1 Long-Horizon Planning and "Compressive Memory" ​

At launch, the official line was that advances in "long-term reasoning and planning" let Devin execute engineering tasks "requiring thousands of decisions." The genuinely valuable technical disclosure came in June 2025, in Cognition engineer Walden Yan's famous blog post "Don't Build Multi-Agents." Counterintuitively, it argues: don't rush into multi-agent architectures, and it offers two context engineering principles:

  1. Share context: share the full agent trace, not just message fragments;
  2. Actions carry implicit decisions: every action embeds implicit decisions, and decision conflicts between parallel subagents are a reliability killer.

Their practical recipe: default to a single-threaded linear agent; when the context window runs out, introduce a model specialized in "history compaction" that compresses actions and conversation history into key decisions and events—Cognition revealed they fine-tuned a small model for this. This maps exactly onto "context is a scarce resource" from the Context Engineering page, and it is important primary material for the reliability trade-offs discussed in Multi-Agent Architecture.

Why Did "Anti-Multi-Agent" Cognition Build Parallel Devins?

No contradiction. What Cognition opposes is multiple agents collaborating on the same task (decision conflicts, fragmented context); Parallel Devins is multiple agents independently completing uncoupled tasks (like assigning work to different remote employees). The former is distributed decision-making; the latter is task-level parallelism. To decide which your scenario needs, ask whether the subtasks must share decision context.

3.2 SWE-bench Score Inflation ​

Putting Devin's score on a timeline makes the capability slope of coding agents over these years vivid:

DateSystemScoreNotes
2023-10Claude 2-era baselines~1.96% (unassisted)SWE-bench paper baseline
2024-03Devin13.86%Random 25% subset of the full set, unassisted
2025-09Claude Sonnet 4.582.0%SWE-bench Verified (human-curated subset)

Note the methodology difference: Devin's 13.86% ran on a random subset of the original SWE-bench, while later high scores mostly ran on the Verified subset—the two are not directly comparable. But the order-of-magnitude change is real: within two years this benchmark went from "double digits makes you a hero" to near saturation, forcing the industry toward longer-horizon new benchmarks (SWE-EVO, Terminal-Bench). This is why you must always ask about methodology before reading any benchmark number—see Agent Evaluation.

3.3 In-House Models: From Consuming to Training ​

Cognition started as a pure "model consumer" and turned to in-house models in 2025: after acquiring Windsurf it inherited the SWE-1 model family (SWE-1, SWE-1-lite, SWE-1-mini, released May 2025, focused on the full software-engineering workflow and long tasks), and in 2026 it shipped SWE-1.6—officially now the most used model inside Windsurf, at up to 950 tok/s. The company also stresses that it is an "independent agent lab": it partners with all foundation model labs and automatically selects the most cost-effective model per task category (they claim to evaluate 100+ categories of software engineering tasks). In-house small models for compaction, third-party flagships for reasoning, and their own models for speed and cost—this hybrid supply strategy deserves study by every agent team that runs the numbers; see Cost Engineering.

4. Business Model and Pricing: From Luxury to Commodity ​

Devin's pricing history is itself a market-education story:

DatePricingContext
2024-03 to 2024-11Waitlist only, targeted invitesManufacturing scarcity while capacity ramped
2024-12$500/month, general availabilityPositioned as an enterprise "AI employee," priced against engineer salaries per head
2025-04 (Devin 2.0)Core $20/month, billed by ACU (Agent Compute Unit), overage ~$2.25/ACU; Team $500/month with 250 ACUs (~$2.00/ACU)Entry price cut 25x
2026Core $20 / Team $500 framework continues; ACU unit price keeps thinning as model costs fallPrice war as the norm

The April 2025 price cut was a defensive move, not a flex. By then Cursor was sweeping individual developers at $20/month, Claude Code had raised terminal-agent experience to a new height, and open-source OpenHands had caught up to the same magnitude on benchmarks. $500/month was entirely unsustainable under a "15% task success rate" reputation. The price cut + ACU metering essentially restructured the business from "hire an AI employee" (per-seat pricing) to "buy compute, pay as you go"—an admission that what it sells is compute, not labor.

This shift is a broadly applicable pricing lesson for the industry: when an agent's success rate can't support pricing on "output," you can only price on "consumption." Outcome-based pricing per issue or per PR appears only once success rates get high enough to promise results—such hybrid models had already appeared in Devin's 2026 enterprise contracts, but usage-based billing remains the backbone.

5. 2025–2026: The Windsurf Acquisition and the Valuation Rocket ​

5.1 The Year's Most Dramatic Acquisition ​

The mid-2025 Windsurf acquisition deserves its own retrospective, because it rewrote the fate of three companies at once:

  1. Spring 2025: OpenAI negotiated an approximately $3 billion acquisition of the AI coding tool Windsurf (parent company Codeium); the talks ultimately collapsed—the Microsoft–OpenAI IP terms were one core obstacle.
  2. July 11, 2025: Google struck lightning, paying roughly $2.4 billion in technology licensing fees plus a talent agreement to bring Windsurf CEO Varun Mohan, co-founder Douglas Chen, and the core R&D team into Google DeepMind. Windsurf instantly became a "shell"—brand, product, customers, and most staff remained, but the soul was gone.
  3. July 14, 2025: three days later, Cognition announced the acquisition of what remained of Windsurf—the product, brand, enterprise customers, and the people who stayed.

The deal is widely viewed as 2025's most complicated AI acquisition: Google took the people and technology licenses, Cognition took the product and revenue, and OpenAI walked away empty-handed. The strategic meaning for Cognition is clear: Windsurf filled in Devin's missing half—the synchronous IDE scenario (Cascade Agent, the SWE-1 model family, a mature enterprise sales pipeline). From then on, Cognition's product matrix became the dual-track structure of "async cloud agent (Devin) + synchronous IDE (Windsurf)," aimed directly at the main battlefield of Cursor and GitHub Copilot.

5.2 Funding and Revenue Curve ​

DateEvent
2024-03$21M Series A (led by Founders Fund)
2024-04$175M at a $2B valuation (reported figures, independently unverified)
2025-08~$500M at a $9.8B valuation
2025-09$400M round
2026-05-27Over $1B at a $26B valuation (led by Lux Capital, General Catalyst, 8VC)

Numbers disclosed alongside the 2026 Series D: annualized revenue of $492 million; enterprise usage up more than 10x since the start of 2026; customers including Citigroup, Mercedes-Benz, Goldman Sachs, Dell, Santander, and the U.S. Army and Navy; Mercedes-Benz compressed a legacy-system modernization project estimated at 8 person-years into 8 days; Latin America's largest bank, Itaú, uses Devin to auto-fix 70% of security vulnerabilities; and 89% of Cognition's own code commits are made by Devin (the rest by Windsurf's local agents).

Maintain professional skepticism when reading these numbers: they all come from the company's own blog, with no third-party audit. But "passing the security and procurement reviews of Goldman Sachs, Citigroup, and the U.S. military" is itself harder proof of capability than any benchmark—those institutions don't pay for viral videos.

5.3 Position Within the Competitive Landscape ​

By 2026 the coding-agent market is stratified. Devin's position and its neighbors:

ProductInteractionAutonomyPricing AnchorTypical User
DevinCloud async, delivered via Slack/PRsHigh (hours unattended)From $20 + ACU meteringEnterprise engineering teams, batch tasks
CursorSynchronous in-IDELow (human watches every change)$20/month per seatIndividual developers
Claude CodeTerminal sync, with cloud tasksMedium (progressive authorization)Subscription + API meteringIndividual engineers and small teams
OpenHandsOpen source, self-hostedHigh (Devin-like)Model API costsTeams that self-host or do research

Devin's differentiation is not "being smarter"—everyone draws on the same model layer—but three things: productized completeness of async long tasks, enterprise compliance and procurement channels (the actual reason military and banking customers choose it), and the "async + sync" dual-track coverage after acquiring Windsurf. Its risk is equally clear: Anthropic's and OpenAI's own agent products are pushing the same capabilities into the same customers at lower prices.

How to Read Vendor-Reported Case Numbers

In comparisons like "8 months compressed to 8 days," the denominator is usually "the most pessimistic estimate of a human team following traditional process," and the numerator usually excludes the human time for requirement clarification, environment prep, and final acceptance. The right takeaway is not the multiplier but the task type: legacy system modernization, bulk security-vulnerability fixes, dependency upgrades—tasks that are "well-defined, high-volume, low per-unit intellectual content" are exactly where autonomous agents shine today. Sending agents to do this is far more reliable than asking them to "develop new features end to end."

6. Controversies and Lessons: Autonomy Hype vs. Reality ​

Devin is the best specimen for studying "agent product marketing"; the lessons distill to four.

6.1 The Phrase "AI Software Engineer" Was Itself the Problem ​

Naming the product "engineer" promised the full job description of an engineer: understanding fuzzy requirements, weighing trade-offs, being accountable for outcomes. What Devin in 2024–2025 actually delivered was "an executor with some success rate on well-defined tasks." The gap between name and capability directly determined the intensity of the backlash—if it had been called "Devin: automated PR bot," there would have been neither the sensation nor the debunking. When naming your agent product, the name is the acceptance criterion users will hold you to.

6.2 Autonomy Is Double-Edged; Correction Points Must Move Earlier ​

The most instructive failure mode in Answer.AI's field test was not "doing it wrong" but "battling an impossible task for a whole day." Full autonomy + long horizons means wrong assumptions compound by the hour. Devin 2.0 moved plan confirmation earlier with Interactive Planning and let humans take over anytime via the IDE—essentially an architectural admission: what today's autonomous agents need is "structured human intervention points," not "less human intervention." This is the core thesis of Human-in-the-Loop design, and the number-one item in Agent Pitfalls.

6.3 A Reputation Trough Isn't the End; Distribution and Trust Rebuilding Are the Long Game ​

By early-2025 discourse, Devin was "already dead": faked demo, botched field test, absurd pricing. But Cognition did three things right: cut the price to a frictionless trial range; entered the much-higher-success-rate scenario of "codebase understanding" with Wiki/Search; and acquired enterprise distribution channels via Windsurf. By 2026 it was one of the highest-revenue agent companies. The lesson is counterintuitive: competition between agent products is not single-task success rate but the engineering of "failure cost × iteration frequency"—whoever drives the cost of failure low enough that users keep trying will be there the day model capability catches up.

6.4 Maintain Institutional Skepticism Toward Benchmarks and Demos ​

After the Devin affair, the industry's immune response was real: SWE-bench officially introduced the human-curated Verified subset and third-party hosted evaluations; vendor-reported numbers now get discounted by default; "releasing full traces" gradually became the minimum bar for a credible release. If you're building your own agent evals, copy this homework: public task sets, public full trajectories, third-party reproducibility—methodology in Evals in Practice.

7. Positive Contributions to the Industry ​

Beyond the criticism, Devin's historical contributions are just as real:

  • It turned "fully autonomous long-task agents" from sci-fi into a category. Before March 2024, the mainstream narrative was Copilot-style completion; after it, "async cloud agents" became every vendor's baseline—GitHub Copilot coding agent, Jules, OpenAI Codex (2025 edition), Claude Code's cloud tasks—all of this category.
  • It educated SWE-bench. The shock of 13.86% turned this academic benchmark into an industry standard, objectively maturing the whole evaluation ecosystem.
  • It pioneered the "sandbox + PR delivery" engineering paradigm. Isolated cloud environments and PR-based integration into human review are now the default architecture for every similar product.
  • It contributed the best architectural polemic. "Don't Build Multi-Agents" remains one of the most-cited engineering blog posts in the context engineering space.
  • It provided a complete business specimen. The pricing correction from $500 to $20, the repositioning from "replacing people" to "engineering leverage," the evolution from product company to "agent lab + model training"—every step is a case study in decision-making.

8. Summary ​

Devin's story is neither an "unmasking of a fraud" nor a "return of the king" saga; it is the agent industry's first complete stress-test report:

  • On the marketing side, it demonstrated the price of overpromising—the capability was real, but once inflated, even the real parts get dismissed as fake;
  • On the product side, it demonstrated the right way to do autonomy—not unsupervised operation, but designing human intervention points into the architecture;
  • On the business side, it demonstrated the strategy for surviving until models catch up—cut prices, switch scenarios, buy channels, then wait.

For learners, the recommended approach: read Answer.AI's field-test original, then Cognition's Series D blog, and compare how the two documents describe the same company—you'll gain sharper judgment than any secondhand commentary can give. To experience a similar architecture hands-on, open-source OpenHands is the most direct alternative; to understand the underlying loop, go back to Agent Loop and Planning.

References ​