Appearance
Case Study: Manus
On March 6, 2025, a Chinese team that got its start building browser extensions for overseas markets — Butterfly Effect, the parent company of Monica — released a four-minute demo video introducing Manus as "the world's first general AI agent": you hand it a one-line goal, and it spins up a computer in the cloud on its own, browses the web, writes code, and crunches data, delivering a website, a report, or a spreadsheet tens of minutes later. The video passed one million views within 20 hours, invitation codes were bid up to tens of thousands of yuan on resale platforms, and within weeks the waitlist topped two million (per multiple media reports).
What makes Manus unusual is not the tasks it demoed — Devin and every Deep Research product were already doing those — but its team makeup and technical stance: this is a company with no in-house foundation model and no plans to build one, which bet everything, from day one, on the system that sits outside the model. Co-founder and chief scientist Yichao 'Peak' Ji put it bluntly in a later engineering blog post:
"If model progress is the rising tide, we want Manus to be the boat, not the pillar stuck to the seabed."
That line is one team's reading of the "Agent = Model + Harness" formula from What Is an Agent Harness: when the underlying model is equally available to every player, competitive advantage can only come from the harness. That is what makes Manus the best specimen for observing the "harness-first" route — its architecture, its published engineering lessons, its business model, even its acquisition and regulatory saga are all footnotes to that route.
Timeline at a Glance
- 2022–2024: Xiao Hong founded Butterfly Effect in Beijing and, after acquiring ChatGPT for Google, launched Monica, an AI browser extension aimed at overseas markets; Yichao 'Peak' Ji, Zhang Tao, and others joined in 2024 (per Sina Finance's account of the team's history, ByteDance offered $30 million for the company in early 2024 and was turned down).
- March 6, 2025: An early preview of Manus went live, invite-only; the team officially claimed state-of-the-art results on the GAIA benchmark. The demo went viral and invite codes got bid up.
- March 28, 2025: Pricing announced — Starter at $39/month (3,900 credits, 2 concurrent tasks) and Pro at $199/month (19,900 credits, 5 concurrent tasks) — along with an iOS app (TechCrunch coverage).
- April 2025: A $75 million round led by Benchmark at a valuation of roughly $500 million (TechCrunch, citing Bloomberg); the company later moved its headquarters to Singapore and wound down its domestic team.
- July 18, 2025: Ji published the engineering blog post Context Engineering for AI Agents: Lessons from Building Manus, laying out six lessons in context engineering (detailed below).
- July 31, 2025: Wide Research launched, able to spin up 100+ full-fledged sub-agents in parallel for a single task (VentureBeat coverage).
- December 2025: The company claimed annualized revenue (ARR) past $100 million within eight months of launch.
- December 29, 2025: Meta announced the acquisition of Manus — reportedly for over $2 billion — with Xiao Hong becoming a VP at Meta while the product continues running its subscription service independently (The Paper's report).
- January–April 2026: Chinese authorities opened a national-security review of the deal; on April 27, 2026, the working mechanism office for foreign-investment security review decided to block the transaction and ordered both parties to restore the status quo ante — the first time the mechanism was publicly used to unwind a completed cross-border deal (MMLC Group's analysis, Trivium China's analysis). Meta reportedly then began isolating Manus from its internal systems and halting the integration (Codersera's follow-up tracking).
Harness Architecture: One Cloud VM per Task
Manus's official self-description is not "an agent" but "a personal cloud computing platform": every session runs on a dedicated cloud VM, where the agent gets a browser, a terminal, code execution, and a filesystem, and the user operates this computer in natural language (company wording quoted in VentureBeat's report).
text
┌────────────── Manus cloud session (one dedicated VM per task) ────────────────┐
│ │
│ ┌────────────┐ ┌────────────┐ ┌────────────┐ ┌────────────────┐ │
│ │ Browser │ │ Term/shell │ │ File system│ │ Code exec / │ │
│ │ research │ │ deps / │ │ read/write │ │ office suite │ │
│ │ web ops │ │ scripts │ │ ext memory │ │ sheets / sites │ │
│ └────┬───────┘ └────┬───────┘ └────┬───────┘ └──────┬───────────┘ │
│ └───────────────┴───┬───────────┴─────────────────┘ │
│ ▼ │
│ ┌────────────────────────────┐ │
│ │ Agent Loop (in-house │ ← reportedly calls Claude │
│ │ framework, rewritten 4× │ and similar frontier models │
│ │ since launch) │ │
│ └─────────────┬──────────────┘ │
│ ▼ │
│ todo.md (the recited plan, │
│ visible live in the user UI) │
└────────────────────────────┬──────────────────────────────────────────────────┘
│ deliverables: website / report / spreadsheet / slides
User: states the goal, walks away, and waits asynchronously for the notificationSet against its peers, Manus's coordinates are clear:
| Product | Task domain | Where it runs | Human's role | Deliverables |
|---|---|---|---|---|
| Manus | General tasks | Cloud VM, async | Assign the job, then wait for the notification | Websites / reports / spreadsheets / slides |
| Devin | Software engineering | Cloud sandbox, async | Assign the job + review PRs | PRs |
| Claude Code | Software engineering | Local terminal, sync | Pairing in the seat, interruptible anytime | Code changes |
| Deep Research tools | Deep research | Cloud, async | Ask a question, then wait | Research reports with citations |
Put the architecture diagram above next to the one in the Devin case study and the two are nearly isomorphic: sandboxed VM + browser + shell + file read/write. The difference is the task domain — Devin narrowed into software engineering (delivering PRs), while Manus expanded into "any general task that can be done on a computer" (delivering websites, reports, spreadsheets, slides). No surprise there: the Manus team has said outright that the maturation of products like Cursor and Devin in 2024 was the direct inspiration for building Manus.
Compared with Devin, Manus made three harness decisions worth calling out on their own:
1. An async-first human interface. The user "states the goal, closes the laptop, waits for the notification." This removes the human from the execution loop entirely, demanding far more autonomy from the harness than Claude Code's "pairing in the seat" model — with no human correcting course mid-run, the agent has to recover from its own mistakes. That is also the product-level reason their "keep the errors in context" lesson (below) matters so much.
2. Trading VM isolation for breadth of task domain. A general agent runs arbitrary scripts, logs into arbitrary websites, and reads and writes arbitrary files — a far larger blast radius than a coding agent working inside one repo. A disposable cloud VM physically contains that blast radius — the same solution Devin uses; see the "sandbox isolation vs. in-process approval" trade-off discussion in Permissions & Human-in-the-Loop.
3. The model is a pluggable part. Manus builds no foundation model of its own; production reportedly runs mainly on third-party frontier models such as Claude, stitched together by an in-house orchestration layer. That means all of its differentiation lives in the harness — which also explains why it was willing to publish its harness lessons in a blog post: what gets shared is "craft," not "assets."
The Six Context Engineering Lessons, Reread Through a Product Lens
That July 2025 engineering blog post is Manus's biggest knowledge contribution to the industry. The Context Engineering chapter already unpacks the technically load-bearing parts one by one (KV cache, mask-don't-remove, the filesystem as context, todo.md recitation, keeping the errors); rather than repeat the technical detail, this section answers a different question from a product and business standpoint: taken together, what do these six lessons say about what kind of company Manus is?
First, all six in full:
- Design around the KV cache: the single most important metric for a production agent is the KV cache hit rate; Manus's average input-to-output token ratio is about 100:1, so whether a token hits the cache is directly a 10x cost difference (Claude charges $0.30 per million cached input tokens vs. $3 uncached).
- Mask, don't remove: tool definitions stay fixed at the top of the context, with logits masking restricting the available actions mid-run, rather than dynamically adding and removing tools — this preserves the cache and keeps the model from hallucinating actions against tool definitions that no longer exist.
- The filesystem is the ultimate context: compression must be recoverable — web content can be dropped as long as the URL survives; document content can be skipped as long as the path still exists in the sandbox.
- Steer attention through recitation: in long tasks averaging about 50 tool calls, repeatedly rewriting todo.md "recites" the global plan toward the tail of the context, countering lost-in-the-middle.
- Keep the errors in context: failed actions and stack traces are the evidence the model uses to correct itself; cleaning up the trajectory amounts to destroying the evidence.
- Don't get trapped by few-shot patterns: highly homogeneous history in the context nudges the model into repetitive inertia (most visible when batch-processing 20 resumes); deliberately inject structured variation to break the pattern.
Why two and a half of the six lessons are about saving money
Note that lessons 1 and 2 are directly about cost, and lesson 3 indirectly so (external storage instead of expensive long-context prefill). That is no coincidence: Manus's business model is a subscription billed in credits — a typical task burns roughly 150 credits, and every credit is backed by real API spend. For a company with its own model, harness optimization pads the gross margin; for a company like Manus that buys third-party models by the token, the KV cache hit rate is the gross margin. Context engineering is not a performance trick here — it is a survival condition for the business model.
Another detail easy for product watchers to miss is this self-deprecating aside: since launch, Manus's agent framework has been rewritten four times, each time because the team found a better way to shape the context; they jokingly call this process of "manual architecture search + prompt tuning + trial and error" "Stochastic Graduate Descent." Four rewrites tell you two things: first, harness design has no theory yet, only experimental science; second, harness iteration (measured in hours) is far faster than model fine-tuning (measured in weeks) — which is precisely the original reason they bet on context engineering instead of building their own model.
Wide Research: 100 "Complete Manus" Agents in Parallel
Wide Research, released on July 31, 2025, is Manus's boldest experiment in the Subagents direction — and a direct challenge to the Deep Research playbook: OpenAI's and Google's Deep Research spend tens of minutes drilling one agent deep into a question; Manus's answer is to go wide — for a task like "compare 100 pairs of running shoes," it instantly spins up 100 concurrent sub-agents, each handling the design, pricing, and stock analysis of one pair, then rolls everything into a sortable table and web page within minutes.
Two architectural details are worth noting:
- Sub-agents are not specialized roles — they are complete Manus instances. There is no "manager agent / retriever agent / writer agent" division-of-labor template; every sub-agent carries the full toolset and can take on any general task independently. This sidesteps the complexity of "designing a role topology per task type," at the cost of every instance dragging along the overhead of the full harness.
- The parallelism rides on a virtualization substrate. In the demo, Ji described Wide Research as "the first application of our optimized virtualization and agent architecture, scaling compute capacity 100x" — in other words, the hard part of Wide Research is not agent intelligence, it is the engineering to spin up and reclaim a hundred VMs at once. More evidence that the VM substrate is this product's real asset.
The public disagreement with Cognition
Read Wide Research alongside Cognition's "Don't Build Multi-Agents" from the Devin case study and you get an industry debate that has not converged: Cognition argues that multi-agent architectures with fragmented context are inherently fragile and prefers a single-threaded linear agent; Manus bet straight down the opposite side with "100 full-featured instances in parallel." The public evidence on both sides is thin — VentureBeat pointed out at the time that Manus offered no benchmark showing parallel beating a single agent running serially. The practical takeaway for readers: what Cognition opposes is task splitting where the contexts never talk to each other, and Manus's parallel tasks (100 pairs of shoes, mutually independent) are exactly the scenario where context does not need to be shared — the two lessons may not conflict, and the boundary of applicability is how decomposable the task is.
Commercialization and the Acquisition: Pricing and Endgame for a Harness Company
Manus's commercial design is itself a textbook in harness economics. It bills by credits: the more complex the task and the more tool calls, the more credits burned — the pricing model directly exposes the cost structure, namely "cost ≈ model API call volume x per-token price," with the harness (KV cache, compression, external storage) as the only lever for cutting costs. Not long after the $39/$199 tiers launched, the lineup was re-sliced into Free / Basic $19 / Plus $39 / Pro $199 (as of July 2025), and credit anxiety — users complaining about "burning credits too fast" — never went away. It is the structural problem shared by every agent product that resells model tokens.
Then comes the most dramatic part of this case. On December 29, 2025 — less than ten months after launch, with a claimed ARR past $100 million — Meta announced it was acquiring Manus, reportedly for over $2 billion, Meta's third-largest acquisition ever. Worth pausing on: what exactly did Meta buy? Not a model (Manus does not have one), not the user relationships (the subscription base is pocket change to Meta), but a team that had taken a general agent harness to a production environment with millions of users, plus that four-times-rewritten body of engineering know-how. The acquisition itself is the priciest valuation to date of the thesis that "a harness is an asset class of its own."
But the story does not stop there. In January 2026, Chinese authorities opened a national-security review of the deal; on April 27, the regulators decided to block it and ordered both sides to restore the status quo ante — the first public use of the foreign-investment security review mechanism to unwind a closed cross-border AI deal. A company that took the "scrub-and-relocate" route overseas (a Chinese team, built on Chinese engineering assets, re-domiciled in Singapore, sold to a US giant) landed squarely in the crossfire of two countries' technology controls. By mid-2026, Meta had reportedly isolated Manus from its internal systems and terminated the integration, while the product continues to operate under its Singapore entity.
Four lessons to take away from Manus
- Harness iteration speed is a strategic asset. Hours-scale context engineering iteration vs. weeks-scale model training iteration determines where a team without its own model should place its chips.
- Every context engineering trick has to translate into money. A 100:1 input-to-output ratio means the KV cache hit rate ≈ gross margin; for products reselling tokens, context engineering is financial engineering.
- The isolation substrate determines the breadth of the task domain. The VM sandbox gave the agent freedom of general action and gave Wide Research the physical basis for a hundred parallel instances — compound interest on an architecture decision.
- A harness is valuable, but a harness team is not necessarily sellable. The $2 billion offer proved the market value of harness engineering; the regulatory veto proved that such assets have risen to the level of national chess pieces. Compliance design for the offshore structure has to move in lockstep with the technical architecture.
Analysis: Did the Harness-First Bet Pay Off?
Back to the underlying question: when models are equally available to everyone, how deep a moat can a harness actually build?
The evidence from Manus cuts both ways. The upside: it really did get from a pure harness to a claimed GAIA SOTA, $100 million ARR in eight months, and a $2 billion offer from Meta — under the model vendors' noses. The downside: the moat was being eroded by two forces at all times — model vendors internalizing harness capabilities (the agent products from Claude and GPT eat the general-task market directly), and harness lessons being publicly copyable (that six-lessons blog post itself accelerates this). Manus's answer was to move the moat from "technique" to "heavy assets": the VM scheduling substrate, behavioral data from millions of users, operational know-how around credit pricing — far harder to copy than prompt tricks.
This matches the dynamic view this site keeps coming back to (see Model vs. Harness): harness and model co-evolve, and no design is ever settle-and-forget; Manus's four rewrites, and its endgame from "general agent" to "absorbed by a giant," are both instances of that pattern. The lesson for practitioners is not "go build a general agent" — it is that harness capability must be consolidated into something a blog post cannot replicate: infrastructure, data, or a closed delivery loop in a vertical domain. Otherwise, when the tide rises, the boat and the pillar go under together.
Further Reading
- What Is an Agent Harness — the conceptual groundwork for the "boat and pillar" metaphor
- Model vs. Harness — attributing competitive advantage when models are commoditized
- Context Engineering — a lesson-by-lesson technical breakdown of the six lessons
- Planning & Task Decomposition — why todo.md recitation works: the mechanics of explicit working memory
- Subagents — the design space: Wide Research vs. the Cognition route
- Permissions & Human-in-the-Loop — the trade-offs of VM sandboxes and async delivery
- Case Study: Devin — the isomorphic architecture in the coding domain, and the opposing multi-agent stance
- Case Study: OpenHands — the open-source counterpart of the same architecture
References
- Manus official blog: Context Engineering for AI Agents: Lessons from Building Manus (2025-07-18, Yichao 'Peak' Ji)
- TechCrunch: Manus launches paid subscriptions and a mobile app (2025-03-31)
- TechCrunch: Manus raises from Benchmark at a ~$500M valuation (2025-04-25, citing Bloomberg)
- VentureBeat: Manus launches Wide Research, spinning up 100 sub-agents in parallel (2025-07-31)
- Sina Finance: Why Meta acquired Manus — team background and startup timeline (2025-12-30)
- The Paper: Meta officially announces Manus acquisition; subscription service to continue (2025-12-30)
- MMLC Group: China Blocks Meta's Acquisition of AI Firm Manus on National Security Grounds (2026-04-29)
- Trivium China: China blocks Meta's Manus acquisition using foreign investment security review (2026-05-18)
- Codersera: Manus AI in 2026 — product and integration outlook after the blocked Meta deal (2026-05-25)