Skip to content

Manus and Agent Applications

At a glance Manus's March 2025 launch ignited a boom in general-purpose agents; this article breaks down the product form, technical foundations, and competitive landscape of autonomous task agents, and examines the controversies around their success rates, costs, and safety.

This page contains time-sensitive material, accurate as of 2025-06; job listings, leaderboards, and product features may have changed since. Verify against the original source before citing.

Manus and Agent Applications ​

1. Background: A "General Agent" Frenzy Ignited by Invite Codes ​

In March 2025, a product called Manus set the AI world on fire, both in China and abroad, without any large-scale marketing. In its demo video, the AI no longer just "chats" — it actually opens web pages, operates software, downloads files, and finally hands you a report a dozen-plus pages long. "From thinking to execution" — that is how Manus positioned itself, and it pushed the concept of the "general agent" in front of the general public for the first time as a complete product. For the full picture of the concept, start with What Is an AI Agent?.

The most visible sign of the frenzy was the invite-code economy: Manus is invite-only, and beta slots were so scarce that "Manus invite codes" were at one point scalped for thousands of yuan on second-hand marketplaces, with dedicated resellers and proxy-buying services springing up. Behind the scarcity lay an extreme supply-demand imbalance, as well as the public's intense imagination about "AI that can actually do work for people" — everyone wanted that admission ticket to a "digital employee".

Manus comes from the Chinese team Monica.im (Wuhan Butterfly Effect Technology, founded by Xiao Hong), previously known for the Monica browser extension. Its sudden popularity forced capital markets to rapidly re-price the agent space: multiple startups announced they were going "All in Agent", big tech rolled out their own plans in parallel, and "Agent engineer" positions began appearing in bulk in the job market. The deeper significance of this wave: it marked the moment agents moved out of the paper-and-demo stage and into productization and commercialization. Viewed on a larger canvas, this is the next milestone after conversational AI in the brief history of AI's evolution — the interaction paradigm is shifting from "I ask, you answer" to "I delegate, you deliver".

2. What Is a General-Purpose Agent Product ​

One-sentence definition: a general-purpose agent product is an AI application that takes natural-language tasks, autonomously plans the steps to execute them, calls external tools, and delivers a complete result. Its essential difference from a chatbot: chatbots "answer", agents "deliver". The former outputs text; the latter outputs outcomes.

A typical task flow can be drawn as a pipeline:

User issues a task (natural language)
      ↓
① Task understanding & planning: break down the goal, draw up an execution plan
      ↓
② Tool calling: browse web pages / write and run code / read and write files / call APIs
      ↓
③ Intermediate feedback & self-correction: retry on errors, adjust the approach
      ↓
④ Deliver the result: report, spreadsheet, presentation, runnable code

To see where the "general agent" sits, compare it with the two other common forms:

DimensionChatbotWorkflow AgentGeneral Agent (e.g. Manus)
Task typeConversation, Q&AFixed automation flowsOpen-ended autonomous tasks
Decision-makingSingle-turn generationPreset nodes, hard-codedDynamic planning, self-correction
Tool useLittle or noneSequenced within the flowFree use of browser / code / files
Adaptability to the environmentNoneAny change outside the flow means failureAdjusts as it executes
DeliverableText answersWorkflow outputComplete artifacts (reports / files / apps)
Typical examplesChatGPTRPA, ZapierManus, OpenAI Operator

A one-sentence litmus test

The essence of a general agent is turning a "prompt" into an "execution plan", and an "execution plan" into "real-world actions". Its evaluation criterion is therefore not "did it say the right thing" but "did it get the job done". This is the most fundamental divide between it and ChatGPT and Conversational AI.

In terms of product form, Manus has one more distinctive trait: cloud-based asynchronous execution. After issuing a task in the conversation, the user can close the page; the agent keeps running in the cloud for minutes or even tens of minutes, and delivers the finished work back through a notification. It feels more like "assigning the AI a job" than "chatting with the AI for a while".

3. Technical Foundations: Four Building Blocks Make an Agent ​

A general agent is not a single model but a system. Taken apart, the foundation consists of four building blocks — remove any one of them and nothing runs:

  1. LLM brain: understands the task, breaks it into steps, and judges intermediate results. Planning ability depends directly on the model's reasoning level — which is one reason agents suddenly "got smarter" after DeepSeek-R1 and Reasoning Models kicked off the "slow thinking" trend: models with stronger reasoning plan more reliably. For the basic mechanism, see Large Language Models (LLM).
  2. Tool calling (Tool Use): the model triggers external tools through structured output (function calling). Manus's capability catalog covers browser operation, code execution, and file read/write; for the principles and their limits, see AI Agents.
  3. Multimodal perception: to "see" web screenshots and "read" charts inside PDFs, an agent needs visual understanding. Manus has built-in multimodal models to handle screenshot-style input; see Multimodal Models.
  4. Sandboxed execution environment: the agent's code runs in an isolated virtual sandbox, which keeps it from polluting the host system and is also the key control point for permissions and security. Private files and account credentials involved in a task can only be accessed within that controlled scope.

The orchestration layer is what strings the four blocks together. The mainstream approach is "loop-style execution" — think → act → observe → think again, until the task is complete or a termination condition is met:

while not task_completed:
    plan = llm.plan(state)          # Brain: what to do next
    result = tools.execute(plan)    # Hands: call the tools
    state.update(plan, result)      # Eyes: observe the results, update state
    if llm.should_stop(state):      # Judge: task done, or needs to ask for help
        break

Take "resume screening" as an example; here is roughly what one real task looks like inside the system:

User task: "Screen 50 resumes and pick the 5 candidates who fit the front-end role"
Step 1  Planning       → download attachments → parse the PDFs → extract skill fields → score against the JD
Step 2  Tool calling   → call the file API to download the resume archive (3 corrupted files found)
Step 3  Self-correction → retry the corrupted downloads; if they still fail, log the reason and skip them
Step 4  Code execution → use Python to compute keyword hit rates and generate a candidate ranking table
Step 5  Delivery       → output a Markdown report + the ranked resume list (with screening rationale)

Note that every step may trigger new LLM calls: understanding files, deciding to retry, writing the analysis script, organizing the report — a task that "looks simple" is backed by dozens of model calls and a dozen-plus tool operations. This is exactly the root of the "runaway cost" controversy, and the reason agent evaluation cannot look only at the "final result" but must also weigh process reliability (methodology in LLM Evaluation and Benchmarks).

Engineering practice at this layer (LangGraph, AutoGPT, the Claude computer-use SDK, etc.) matters enormously: the framework determines an agent's robustness, cost, and observability. To build a minimal working agent with your own hands, follow Building an Agent from Scratch through the entire process. For the failure modes and anti-patterns that show up in practice, cross-check Common Pitfalls and Anti-Patterns.

An easily overlooked cost item

A single agent run often invokes the LLM dozens to hundreds of times, and multimodal inputs further amplify token consumption. For the same task, the agent path can cost tens of times more than a direct conversation. Cost modeling and inference optimization are the lifeblood of agent commercialization; for the methods, see Inference Optimization and Quantization.

4. The Product Landscape (dataAsOf: 2025-06) ​

Manus is not an isolated case. From 2024 to 2025, AI products that "can actually do work" flourished, each with its own emphasis:

ProductVendorPositioningHighlights
ManusMonica / Butterfly Effect (China)General autonomous task agentCloud async execution, complete deliverables, multi-model routing
OpenAI OperatorOpenAIBrowser-operating agentBuilt on the CUA model; operates web pages to order food, shop, and more
Claude Computer UseAnthropicComputer-operating APIDirectly controls the screen cursor and keyboard; open to developers
DevinCognition AICoding agentEnd-to-end dev tasks: create repos, fix bugs, deploy
Genspark AutopilotGensparkAI search + agentUpgraded from "searching for answers" to "getting things done for you"
OpenAI Deep ResearchOpenAIResearch agentMulti-step web research; outputs cited research reports

How to read this landscape

The differences between these products fall mainly along two axes: permission boundary and domain focus. Operator and Computer Use control "the browser/computer"; Devin focuses on "the codebase"; Deep Research specializes in "research"; Manus aims to "do everything". The broader the permissions, the higher the capability ceiling — and the greater the risk and cost. That rule holds for every agent product.

In terms of technical lineage, coding agents continue the line of GitHub Copilot and Code Intelligence, but extend the scenario from "completing code" to "independently completing engineering tasks", making it one of the most mature verticals for agentification.

5. Use Cases and the Limits of Productivity Agents ​

Judging from Manus's official demos and hands-on user reports, high-frequency scenarios cluster into the following categories:

  • Resume screening: batch-parse PDF resumes, score and rank them against JD keywords, and output a candidate comparison table;
  • Web research: given a topic, automatically visit multiple sites, gather information, and compile a sourced research report;
  • Data analysis: read CSV/Excel files, write Python code for cleaning and statistics, and output charts and conclusions;
  • Report generation: one-stop report production, from gathering material to a typeset document;
  • File organization: download, classify, and rename files, and maintain a local knowledge base.

What these scenarios share is: a clear goal, a decomposable process, and verifiable results. The flip side is that the boundaries are just as clear:

The limits of productivity agents

Open-ended, ambiguous tasks that depend on human judgment (such as "help me decide whether this company is worth investing in") are not suitable for agents; tasks involving private or high-value data demand extreme caution; tasks with stringent real-time requirements are constrained by execution speed and the tool ecosystem. The test in one sentence: only hand over what can be verified. For any task whose acceptance criteria cannot be spelled out, the agent will simply run further and further in the wrong direction.

6. Controversies and Risks ​

Beneath the hype, the criticism is just as sharp, concentrated on four dimensions:

  • Task success rate in doubt: demo videos are edited, and the gap between "demo vs. real-world usable" objectively exists. In community testing, complex tasks frequently stall midway, get half the job wrong, or drift completely off track; the real success rate is nowhere near 100%, especially for complex, long-running tasks.
  • Runaway costs: each task consumes large amounts of tokens (planning + multi-turn tool calling + multimodal inputs), compounded by cloud async execution — a single task can cost tens of times more than a traditional conversation. Pricing and margins are unavoidable problems for agent productization: "free flashy demos" are easy; "profitable delivery" is hard.
  • Permissions and security: an agent operates the browser and reads and writes files using your accounts. If it runs into prompt injection (a malicious web page tricking the agent into dangerous actions), the damage ranges from leaked information to actual losses. For the overall security governance framework, see AI Safety and Governance.
  • Missing benchmarks: the open-endedness of agent tasks renders traditional LLM benchmarks like MMLU ineffective; agent benchmarks such as GAIA have limited coverage, and a gap remains between "great exam takers" and "agents that actually get work done". For how to evaluate an agent systematically, see LLM Evaluation and Benchmarks.

Permissions are an agent's first security boundary

Permissions granted to an agent should follow the principle of least privilege: grant only the tool scopes and data scopes the task actually requires, and make human confirmation mandatory for critical operations. "The stronger the capability, the more restrained the use" — the quality of permission design often decides a product's fate more than model selection does.

7. Industry Impact: From "Conversation" to "Delivery" ​

The agent boom is reshaping the AI industry's narrative on three levels:

  1. Agent as a Service (AaaS): business models evolve from "selling API calls" and "selling conversation quotas" to "charging per completed task". What users buy shifts from "compute" to "results" — a fundamental shift in the value anchor.
  2. Paradigm shift: the way people use AI moves from "I ask, you answer" to "I delegate, you deliver". This transforms an entire chain — product interaction, evaluation methods, billing rules, and security systems — and turns "capability boundaries" into a piece of fine print every product must spell out. For the panoramic context, see What Is AI? The Hot Concepts and The Concept Landscape.
  3. Talent structure shifts: the core skill of agent developers moves from "tuning prompts" to "designing tools, orchestrating workflows, and controlling cost and risk". For job seekers, agent engineering has become a high-value differentiator; for a breakdown of the roles and skills, see JD Knowledge Map.

Compared with the "conversation revolution" sparked by ChatGPT and Conversational AI, the agent wave Manus represents is an "execution revolution": the former changed how humans obtain information; the latter attempts to change how humans complete tasks. This road has only just begun, but the direction is already clear — in the second half of AI, the game moves from "being able to say" to "being able to do". And "doing" implies responsibility and boundaries — precisely the question everyone in the field must answer together in the stage ahead.

Further Reading ​

References ​