Skip to content

Manus

At a glance The general-purpose AI agent that broke out in March 2025 on the strength of one demo video: invite codes scalped for tens of thousands of dollars, self-reported SOTA across all three GAIA tiers, then a $2 billion sale to Meta a year later that Chinese regulators blocked. This page dissects its multi-agent architecture, VM sandbox, and that Context Engineering blog post the whole industry forwarded.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Manus ​

Manus is a general-purpose AI agent product released in March 2025 by the Chinese team Butterfly Effect (maker of Monica). The name comes from the Latin word for "hand," signifying the extension of AI from "thinking" to "doing." It was the first product to truly break out under the "general agent" banner: on launch day the whole internet hunted for invite codes, resale platforms scalped them for up to tens of thousands of yuan, and a batch of "Manus concept stocks" on the A-share market surged to their daily limits.

For learners, Manus's value lies beyond the product itself. Its later trajectory—viral fame, fundraising, relocation to Singapore, a $2 billion sale to Meta, blocked by Chinese regulators, independent operations restored in August 2026—is practically a condensed history of the general-agent space. And the Context Engineering blog post the Manus team published in July 2025 remains one of the most valuable public engineering documents for building production-grade agents.

1. A Phenomenon Launch: The March 2025 Breakout ​

Launch and Spread ​

In the early hours of March 6, 2025, Monica.im launched Manus with a demo video, calling it "the world's first general-purpose AI agent." Founder Xiao Hong is a post-90s serial entrepreneur (previously behind the Monica browser extension); co-founder and chief scientist Yichao "Peak" Ji is an NLP veteran. In the demo, Manus showed screening resumes, deep stock analysis, real-estate research, website generation, and more. The difference from chat-style AI: it delivers task results directly instead of just offering advice.

The spread exceeded everyone's expectations:

  • The beta used an invite system, and the website crashed repeatedly under traffic surges;
  • Invite codes were scalped on resale platforms for anywhere from 999 to 50,000 yuan, with isolated listings reported at 88,000 or even 100,000 yuan;
  • Weibo trending topics and marketing accounts everywhere; a batch of A-share companies with little to do with AI were labeled "Manus concept stocks" and rallied;
  • Word of mouth split quickly among users who got codes: some marveled at the automation, while others hit lag, slowness, and unfinished tasks.

Hunger Marketing Is a Double-Edged Sword

The invite system + demo video combo created scarcity and a viral spark, but it also drew down trust. When real users found the experience fell short of the demo, the backlash of "over-marketing" and "hype" followed. For every agent product launch since, "will you open it to field testing?" became almost the media's first question.

The GAIA Claims: Read with a Discount ​

Manus's website claimed SOTA on all three difficulty tiers of the GAIA benchmark (General AI Assistants, which measures general AI assistants' ability to solve real-world problems), with comparison numbers against OpenAI Deep Research:

GAIA TierManus (self-reported)OpenAI Deep ResearchPrior SOTA
Level 1 (basic tasks)86.5%74.3%67.9%
Level 2 (multi-step reasoning)70.1%69.1%67.4%
Level 3 (complex orchestration)57.7%47.6%42.3%

But there is one key detail: at launch (leaderboard data with a March 5, 2025 cutoff), the official GAIA leaderboard had no evaluation record for Manus—these numbers were self-tested and self-reported. Self-reported SOTA is not rare in the agent industry (eval config, pass@1 methodology, and tool budgets all affect results); when reading any agent product's benchmark claims, first check whether third parties have reproduced them. More on evaluation methodology: Evaluation & Evals.

2. Positioning: General-Purpose Task Execution ​

Manus's positioning in one sentence: give AI a cloud computer and let it do the work for you. It is not a vertical tool (coding only, search only) but a general-purpose execution layer. Typical task types include:

  • Deep research: cross-site searching, reading, and synthesis, outputting cited research reports (benchmarking OpenAI Deep Research);
  • Data analysis: scraping data, cleaning it, analyzing with Python, generating charts and reports;
  • Website and app generation: generating and deploying an accessible website straight from a requirements description;
  • Office automation: screening resumes, building slide decks, organizing spreadsheets, price comparison, booking travel.

A few notable design choices in the product shape:

  • Cloud async execution: tasks run on Manus's cloud VMs and keep going after the user closes the browser, with a notification on completion. This is a fundamentally different product philosophy from local agents (like Claude Code).
  • Visible process: users can watch a live replay of the agent's operations—which web page it opened, what command it ran, which file it's writing. This "watching AI work" transparency is the core reason its demos spread so well.
  • Deliverable-oriented: the final output is files (reports, spreadsheets, websites), not a conversation.

Commercially, Manus uses subscriptions with tiers from $19 to $199 per month (per August 2025 disclosures), and free accounts get a limited basic quota. Wide Research, launched in August 2025, allows dispatching hundreds of sub-agents in parallel for large-scale research tasks and is included in the $199/month top tier.

What a Typical Manus Task Looks Like ​

Take "analyze Tesla's latest earnings report and produce a research report with charts" as an example. The actual execution flow looks roughly like this:

  1. Planning: the planning module splits the request into sub-goals and writes them into todo.md (fetch the earnings report → extract key financials → plot with Python → write the report → verify facts and citations).
  2. Retrieval and browsing: the agent drives a headless browser in the sandbox to visit earnings pages, news sources, and third-party data sources, switching sources on its own when it hits login walls or anti-scraping.
  3. Data processing: scraped data is saved as files in the sandbox; Python scripts clean it, compute YoY/QoQ, and generate charts—when a script errors, it reads the stack trace, fixes the code, and reruns by itself.
  4. Writing and assembly: the report is written based on the intermediate artifacts in the sandbox, with charts referenced as files.
  5. Delivery: the output is a packaged bundle (Markdown/HTML report + charts + data tables), with the full operation replay preserved for user auditing.

Throughout, users can interject at any time to correct course—when new input arrives, Manus replies to acknowledge it immediately instead of grinding on (this is exactly the design mentioned in the "tool masking" section below). According to estimates by practitioners involved in early field tests, a single research-type task costs on the order of $2 in inference—about a tenth of OpenAI Deep Research at the time (from podcast discussions, not official disclosure; verify before citing).

3. Technical Architecture: A Puzzle from Public Information ​

Manus is not open source; its architecture information comes from official introductions, team interviews, and third-party teardowns. The assembled picture looks roughly like this:

┌──────────────────────── Manus cloud ────────────────────────┐
│                                                             │
│  User request ──> ┌────────────┐                            │
│                   │  Planning  │ breaks the task into       │
│                   │  module    │ sub-goals / todo.md        │
│                   └─────┬──────┘                            │
│                         ▼                                   │
│         ┌─────────────────────┐   Tool set (~29 tools)      │
│         │  Execution agent    │ ──> Browser (browser_*)    │
│         │  loop (multi-agent  │ ──> Shell (shell_*)        │
│         │  coordination:      │ ──> File read/write /      │
│         │  plan/execute/      │     code execution         │
│         │  verify)            │                            │
│         └────────┬────────────┘                            │
│                  ▼                                          │
│      ┌────────────────────────┐                             │
│      │  Dedicated VM sandbox  │  One "cloud computer" per   │
│      │  (Linux env + file     │  task: install software,    │
│      │  system)               │  run scripts, store         │
│      └────────────────────────┘  intermediate artifacts      │
│                                                             │
│  Underlying models: Claude (primary) + Qwen-based           │
│  fine-tuned models (not a proprietary foundation model)     │
└─────────────────────────────────────────────────────────────┘

Key points:

  • Multi-agent architecture: officially, complex tasks are decomposed into planning, execution, and verification submodules, each backed by its own model and coordinated via APIs. This matches the mainstream form discussed in the Multi-Agent Architecture chapter.
  • VM sandbox: each task runs in a dedicated cloud Linux VM, operating like Anthropic's Computer Use—the agent owns a full operating system, can install dependencies, run code, and operate a browser, rather than being limited to a few pre-wrapped APIs. Per official figures disclosed in December 2025, Manus has cumulatively created more than 80 million virtual computers.
  • No foundation-model training: the underlying layer relies on frontier models like Claude plus Qwen-based fine-tuned models (on March 11, 2025, Manus announced a strategic partnership with Alibaba's Qwen team). This was a deliberate strategic choice—see below.

The Context Engineering Blog: Manus's Most Important Technical Legacy ​

On July 18, 2025, Yichao "Peak" Ji published "Context Engineering for AI Agents: Lessons from Building Manus," making Manus's core engineering experience public. The article opens with the thesis: the team once faced a choice between "training an end-to-end agentic model on top of open source" and "doing engineering on top of a frontier model's in-context learning," and chose the latter—"if model progress is a rising tide, we want to build a boat, not a post nailed to the seabed." The whole agent framework was rewritten four times for this; the team jokingly calls the process "Stochastic Graduate Descent."

The article's six lessons were all bought in production; here is the distilled essence (a more systematic treatment: Context Engineering):

1. Design around the KV-cache Ji considers the KV-cache hit rate the single most important metric for a production agent. The input/output ratio of an agent is extremely lopsided (Manus averages about 100:1—the context keeps growing while each step's output is just a short function call), and whether the cache hits directly determines latency and cost: for Claude Sonnet, cached input tokens cost $0.30 per million versus $3 per million uncached—a 10x difference. The practical requirements: keep the prompt prefix stable (don't put a to-the-second timestamp at the start of the system prompt), make context append-only, keep JSON serialization deterministic, and mark cache breakpoints explicitly when needed. The cost implications of this lesson are expanded in the Cost Optimization chapter.

2. Mask, don't remove When the tool count grows, the intuitive move is to dynamically load tools with RAG, but Manus's experimental conclusion: do not add or remove tool definitions mid-iteration unless absolutely necessary—tool definitions sit at the front of the context, so any change invalidates the entire downstream KV cache; and historical actions referencing deleted tool names confuse the model. Manus's solution is a context-aware state machine managing tool availability, masking disallowed token logits at decoding time while the tool definitions themselves stay untouched. Using response prefill, it can also implement three function-calling constraint modes (auto / required / specified). Combined with deliberate naming conventions (browser tools all prefixed browser_, shell tools all prefixed shell_), you can constrain the action space by group without writing a stateful logits processor.

3. Use the file system as context A 128K context window is insufficient—or even a burden—in real agentic scenarios: observations like web pages/PDFs are often overlong, long-context performance degrades, and transmission and prefill both burn money. Manus treats the file system as an unlimited, persistent, agent-readable-and-writable "external memory." All compaction strategies are designed to be restorable: web content can be dropped from the context as long as the URL is kept; document content can be omitted as long as the path remains in the sandbox. The logic is hard: the agent must predict the next step based on the full historical state, you cannot know which observation becomes critical ten steps later, so any irreversible compression carries risk.

4. Manipulate attention through recitation Anyone who has used Manus notices it loves creating a todo.md and checking items off when handling complex tasks. That's not cuteness—it's deliberate attention manipulation: Manus averages about 50 tool calls per task, and in long loops models easily drift and forget the original goal. Constantly rewriting the todo list effectively "recites" the global plan to the end of the context repeatedly, pushing it into the model's recent attention range and sidestepping lost-in-the-middle—no architecture change, just bending the model's focus back to the task goal in natural language.

5. Keep the wrong stuff in In multi-step tasks, failure is the norm, not the exception. The common impulse is to erase traces of failure, retry, or reset state, but "erasing failure is destroying evidence." Keeping failed actions and their errors and stack traces in the context lets the model implicitly update its internal beliefs and lowers the probability of repeating the same mistake. Ji argues that error recovery is one of the clearest markers of genuinely agentic behavior—and it is precisely the part academic benchmarks cover least.

6. Don't get few-shotted Models are excellent imitators: when the context is full of similar action-observation pairs, they repeat the pattern on inertia even when it's no longer optimal. On tasks like batch-processing 20 resumes, the agent falls into a fixed rhythm, causing drift, overgeneralization, even hallucination. Manus's countermeasure is injecting a little structured variation into the serialization templates, wording, order, and format of actions and observations—controlled randomness to break the pattern.

Turning the Six Lessons into Code ​

The following skeleton code demonstrates how to implement four of the lessons at once (append-only context, tool masking, todo recitation, error retention) in your own agent loop; use it directly as an experiment starting point:

python
# A minimal agent loop skeleton embodying Manus's lessons
# (pseudocode; can be wired to any chat API)
SYSTEM_PROMPT = "You are a general-purpose task execution agent."  # Note: no timestamp; stable prefix to hit the KV cache

# Tool definitions go into the front of the context once and never change — protects the KV cache
TOOLS = [browser_open, browser_click, shell_exec, file_read, file_write]  # unified naming prefixes

def agent_loop(task: str):
    context = [system(SYSTEM_PROMPT), user(task)]   # append-only: only append, never modify history
    write_file("todo.md", f"# Task goal\n{task}\n- [ ] in progress")  # recitation: write the goal into a file

    step = 0
    while True:
        # Re-read todo.md into the end of the context every 5 steps to fight lost-in-the-middle
        if step % 5 == 0:
            context.append(tool_result("file_read", read_file("todo.md")))

        # Tool masking: definitions untouched; only constrain the allowed actions by current state
        allowed = state_machine.allowed_tools()      # e.g., when a user interjects, only "reply" is allowed
        action = model.chat(context, tools=TOOLS,
                            allowed_tools=allowed)   # implemented via logits mask / prefill underneath

        try:
            observation = execute(action)            # execute in the sandbox
        except Exception as e:
            observation = format_traceback(e)        # error retention: full stack trace also goes into the context
        context.append(assistant(action))
        context.append(tool_result(action.name, observation))
        step += 1

        if action.name == "finish":
            break

def compress(context):
    # Restorable compaction: drop the body of long observations but keep the URL / file path
    return [drop_body_keep_ref(m) if is_huge(m) else m for m in context]

Note what this code deliberately does not do: no del TOOLS[i] mid-run, no reordering of history messages, no context reset after failure. These three are exactly the taboos beginners most often violate—and that Manus learned the hard way, with real money.

Why This Post Deserves a Close Read

These six lessons cover the three most expensive classes of problems in agent engineering: cost and latency (KV cache), behavioral stability (tool masking, anti-few-shot rigidity), and long-task reliability (file system as external memory, recitation, error recovery). It ranks alongside Anthropic's "Effective Agents" and "How we built our multi-agent research system" as the most-cited practice literature in agent engineering in 2025. Read the original; this page is only a distillation.

4. Business and Capital Timeline (as of August 2026) ​

Manus's capital story is more dramatic than its product story. The complete timeline:

DateEvent
2023-02 / 2023-08ZhenFund seed round (~$14M post-money) and angel round (~$50M post-money)
2024-11Series A with Sequoia China, Tencent, ZhenFund, Wang Huiwen and others participating; ~$85M post-money valuation
2025-03-06Manus launches; viral overnight
2025-04$75M Series B led by Benchmark; valuation near $500M post-money, up ~5x in about six months
2025-06HQ relocates to Singapore; China team cut from ~120 to ~40; website blocks Chinese IPs
2025-08Disclosed annualized revenue run-rate of ~$90M; Benchmark's investment reviewed by the U.S. Treasury under outbound AI investment restriction rules
Early December 2025Official disclosure: over 147 trillion tokens processed cumulatively, over 80 million virtual computers created; ARR passed $100M, claimed as the world's fastest startup from zero to $100M ARR
2025-12-29/30Meta announces the acquisition of Manus's parent Butterfly Effect for over $2 billion; founder Xiao Hong was slated to become a Meta VP
2026-01-08China's Ministry of Commerce steps in to review technology export compliance
2026-04-28The NDRC announces the acquisition is halted
2026-06Bloomberg reports Meta has completed operational separation from Manus and stopped data sharing; the acquisition moves toward unwinding
2026-08-11Manus announces it will resume operating as an independent company, with Tencent and other original shareholders leading a buyback from the Chinese consortium; some user data generated on and after the acquisition day is slated for deletion/migration

A few notable points:

  • The Singapore relocation controversy: taking U.S. venture money, using U.S. model APIs, blocking Chinese IPs, cutting the domestic team—this series of "severing" moves triggered an enormous backlash in mid-2025 and foreshadowed the later regulatory scrutiny.
  • Squeezed from both regulatory directions: the U.S. reviewed Benchmark's investment in Manus (worried about touching outbound AI investment restrictions on China), while China halted Meta's acquisition. An application-layer agent company caught in the crosshairs of both great powers' regulators is a first for the AI industry.
  • The outcome: after 16 months of turmoil, back to square one—independent operation, Chinese shareholders (a Tencent-led consortium) back in control, the valuation anchor still around $2 billion, but Xiao Hong's Meta VP appointment is now moot.

A Telling Contrast

Manus proved that "don't train models, just build the application layer" can produce the fastest commercialization speed (8 months to $100M ARR), but its geopolitical ordeal also shows: application-layer companies are not neutral territory—which models you use, whose money you take, and where your data lives all become regulatory variables. That is a real constraint any AI startup going global in 2026 must factor in.

5. The Controversy: The "Wrapper" Debate ​

Manus has carried the "wrapper" accusation since day one—arguably one of the biggest flame wars in the Chinese AI scene in the first half of 2025.

The critics' view: Manus has no proprietary foundation model; its core capability comes from third-party APIs like Claude, making it essentially an "integration operating system"; the invite system is hunger marketing; the all-English website and overseas payment requirements contradict the "Chinese team" narrative; and developers used the open-source project OpenManus (from the MetaGPT team) to replicate part of its capability in short order, suggesting the moat is shallow.

The defense: making task decomposition, tool orchestration, sandbox environments, and context management usable and stable is itself engineering capability with an extremely high bar. "Building foundation models and building applications are two different trades"; the application layer's value shouldn't be dismissed with the word "wrapper." Ji later also publicly explained that the product uses Qwen-based fine-tuned models, not simple API forwarding.

My judgment, in three layers:

  1. "Wrapper" is factually true and commercially not an insult. The Manus team has never been coy about it—the Context Engineering blog opens with the "boat on the tide" choice. The real question: will model vendors eat the application layer while they're at it? OpenAI's Operator and Deep Research and Anthropic's Claude agent capabilities moved fast in 2025, showing the threat is real.
  2. Part of the moat's answer hides in engineering details. OpenManus can replicate the interface and the flow, but not the KV-cache hit rate, the tool-masking state machine, or the restorable compaction—engineering parameters tuned at the scale of millions of users. The Context Engineering blog's enormous influence is precisely because it showed the thickness of the "shell."
  3. But a pure application-layer moat is indeed limited. Manus's eventual sale to Meta (and later return to Chinese capital) partly confirms the survival pressure on independent general-agent products caught between giants. Compare Devin's deep cultivation of the vertical coding scenario; differentiation is harder to build for general agents.

The debate's industry value: it put the contradiction of "model capability vs engineering value" on the table; in the year after, nearly every agent startup's fundraising deck had a page answering "what is your moat."

6. What Manus Meant for the General-Agent Space ​

Setting the controversy aside, Manus left several indelible markers on the industry's history:

  1. It defined the product paradigm of the general agent. The combination "cloud VM sandbox + tool calling + async execution + process replay + file deliverables" became the standard configuration of general-agent products afterward. Look at the products today: Manus's shadow is everywhere.
  2. It delivered the first mass-market education about agents. Before Manus, "AI agent" was an abstraction to most people; with the visceral experience of "watching AI do the work for you," it made the shift "from chat to work" sink in. If the definition in What Is an AI Agent feels abstract, Manus's demo video is the best illustration.
  3. It contributed one of the most valuable engineering methodologies. The Context Engineering blog pushed "context engineering" into standard industry vocabulary, and its lessons (KV cache first, tool masking, file system as context, error retention) are now required reading for Agent Loop design.
  4. It provided a textbook launch-marketing case, in both directions. The invite-system hunger marketing's viral efficiency and its trust backlash were equally striking; successors like Manus's peers and open-source replicas learned from both.
  5. It became the first complete specimen of an AI application company under geopolitical contest. Starting as a Chinese team, taking U.S. funds, relocating to Singapore, acquired by an American giant, blocked by Chinese regulators, returning to Chinese capital—this 16-month curve is an unavoidable case for understanding the AI startup environment of 2026.

For hands-on learners, the most direct takeaway: read the six Context Engineering lessons thoroughly, then validate them one by one in the practice of building an agent yourself—they are closer to production's real problems than any framework tutorial.

References ​