Appearance
AutoGen
1. What It Is — and the 2026 Reality You Must Know First
AutoGen is the multi-agent framework open-sourced by Microsoft Research, with the paper posted and code released in August 2023 (arXiv:2308.08155). It founded the "conversational multi-agent" school: multiple agents with distinct roles collaborate on a task through natural-language conversation, rather than being driven by a predefined state graph. Its core abstractions influenced nearly every multi-agent framework that followed, and the "two agents talk their way to a result" paradigm went viral starting here.
But discussing AutoGen in August 2026 requires putting one overwhelming fact on the table first:
AutoGen has entered maintenance mode
In October 2025, Microsoft announced that AutoGen and Semantic Kernel would merge into the unified Microsoft Agent Framework (MAF); MAF shipped its 1.0 GA in April 2026. The AutoGen repo's README now states plainly: no new features will be accepted, the project is community-maintained, new projects should use MAF directly, and existing projects should migrate per the official migration guide.
That doesn't make this page worthless, for three reasons:
- A massive amount of existing code and tutorials. Multi-agent tutorials, paper replications, and interview questions from 2023-2025 are full of AutoGen APIs — you will need to read them sooner or later.
- The concepts are still universal. Team, termination conditions, speaker selection — MAF and other frameworks inherited these concepts wholesale, so learning AutoGen means learning the intellectual source of the whole school.
- Interviewers love it. Multi-agent orchestration is a high-frequency topic in agent job descriptions, and "what's the difference between AutoGen's GroupChat and LangGraph's graph orchestration" is a classic question.
The one-line positioning: study it as the living fossil and idea library of multi-agent architecture — don't pick it as a production framework for new projects in 2026.
2. Version History: The 0.2-to-0.4 Ground-Up Rewrite
AutoGen's history splits cleanly into two eras:
The 0.2 era: the classic conversation API (2023-2024)
The old API took ConversableAgent as its base class; typical code looked like this:
python
# old 0.2 API (deprecated; shown only so you can recognize old tutorials)
from autogen import AssistantAgent, UserProxyAgent
assistant = AssistantAgent("assistant", llm_config={...})
user_proxy = UserProxyAgent("user_proxy", code_execution_config={...})
user_proxy.initiate_chat(assistant, message="Draw me a stock price chart")AssistantAgent proposes ideas, UserProxyAgent stands in for the human and executes code, and initiate_chat steps on the gas to start the conversation; multi-party scenarios use GroupChat + GroupChatManager. This API was extremely easy to grasp and single-handedly popularized the multi-agent concept, but the engineering problems piled up: synchronous blocking, messy typing, hidden global state, painful debugging, and no fine-grained control over message flow.
How to tell old tutorials from new at a glance
from autogen import ... or initiate_chat(...) means the old 0.2 API, deprecated since 2025; the new API imports entirely from the three packages autogen_agentchat, autogen_core, and autogen_ext. Most tutorials still online are the old kind — copying them will trip you up.
The 0.4 era: the full rewrite (released January 2025)
In January 2025, Microsoft shipped AutoGen 0.4 — officially described as a rewrite from scratch. The core changes:
| Dimension | 0.2 | 0.4+ |
|---|---|---|
| Execution model | Synchronous, blocking conversation loop | Fully async, event-driven (async/await) |
| Packaging | Single package pyautogen | Split into autogen-core / autogen-agentchat / autogen-ext |
| Architecture | Flat; agents call each other directly | Layered: Core runtime → AgentChat high-level API → Extensions |
| Observability | Essentially none | Built-in tracing, message streaming, OpenTelemetry |
| Cross-language | Python only | The Core layer is designed for a cross-language .NET/Python runtime |
| Typing | Weak | Fully type-annotated |
The layered architecture can be understood like this:
┌───────────────────────────────────────────────────────────────┐
│ App layer AutoGen Studio (no-code GUI) · Magentic-One │
├───────────────────────────────────────────────────────────────┤
│ AgentChat AssistantAgent · Teams · termination conditions│
│ ← the only layer most people ever need │
├───────────────────────────────────────────────────────────────┤
│ Core event-driven actor runtime · messaging │
│ ← drop down here only for fine-grained control│
└───────────────────────────────────────────────────────────────┘
Extensions cut across every layer: model clients / code execution / MCP tools- Core: an event-driven actor runtime where agents communicate only through async messages; supports local and distributed deployment. Flexible but verbose — the layer for framework-level customization.
- AgentChat: the "opinionated" high-level API on top of Core — preset agents and preset teams, the first choice for quick prototypes, and the closest to the 0.2 mental model.
- Extensions: pluggable implementations — model clients (OpenAI, Azure OpenAI, etc.), the Docker code executor, MCP workbenches, and more.
Iteration continued after 0.4; the last stable minor line before maintenance mode was 0.7.x (autogen-agentchat 0.7.5, released September 2025), requiring Python ≥ 3.10. No feature-bearing major version followed.
3. Core Concepts: Agents, Teams, and Termination
The AgentChat layer in 0.4+ has only three core concepts; grasp them and you've grasped the framework.
AssistantAgent and the built-in agents
AssistantAgent is the workhorse: it wraps an LLM and can carry tools, a system message, and memory. Other built-ins include UserProxyAgent (human in the loop — waits for human input each round), CodeExecutorAgent (executes code and returns the results), and MultimodalWebSurfer (browser operation). Tools are ordinary Python functions whose schemas are generated automatically from type annotations and docstrings; you can also attach MCP servers directly (see Tools & MCP).
Teams: four preset orchestration modes
A Team is a group of agents plus a set of rules governing "who speaks next and when to stop." AgentChat ships four presets:
| Team | Speaker-selection mechanism | Best for |
|---|---|---|
RoundRobinGroupChat | Fixed order, taking turns | Deterministic flows like two-agent reflection (writer + critic) |
SelectorGroupChat | An LLM picks the next speaker each round | Free-form discussion with many roles and no fixed flow |
MagenticOneGroupChat | An Orchestrator assigns work dynamically via a ledger | Open-ended web/file/code tasks (next section) |
Swarm | Agents hand off proactively via HandoffMessage | Support-style "transfer" flows |
RoundRobinGroupChat is cheap and predictable — the default starting point. SelectorGroupChat is flexible but adds one extra LLM call per round just to choose the speaker, so cost and latency climb noticeably as rounds accumulate.
Single agent first, team later
The official AutoGen docs themselves urge this: teams need more steering scaffolding, so optimize a single agent's tools and instructions first, and move to a team only once you've proven the single agent isn't enough. This matches the conclusion of the Multi-Agent Architecture page — for most tasks multi-agent is over-engineering, and a writer-critic pair already captures most of the real benefit.
Termination conditions: the gate that keeps conversations from burning your budget
The biggest engineering risk in conversational orchestration is "talking forever" — agents being polite to each other, correcting each other, the token bill climbing exponentially. AutoGen makes termination conditions a first-class citizen, combinable with bitwise operators:
python
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination
# Stop when the critic says APPROVE; but force a stop after 20 messages no matter what (the backstop)
termination = TextMentionTermination("APPROVE") | MaxMessageTermination(20)Also common: TimeoutTermination (by time), TokenUsageTermination (by token usage), and ExternalTermination (stopped by an external signal). The iron rule of engineering practice: a semantic condition (such as a keyword) plus a hard cap (message count / tokens) always appear as a pair. This is also the most basic line of defense for cost and budget control in multi-agent scenarios.
4. Magentic-One: the Blueprint for a Generalist Multi-Agent System
In November 2024, Microsoft Research's AI Frontiers lab released Magentic-One (paper arXiv:2411.04468), built on AutoGen and positioned as a "generalist" multi-agent system: instead of tuning prompts for specific tasks, a lead Orchestrator coordinates four specialist agents to complete open-ended web, file, and code tasks.
┌─────────────────┐
│ Orchestrator │ Outer loop: Task Ledger (known facts + the plan)
│ (lead agent) │ Inner loop: Progress Ledger (progress + next step)
└────────┬────────┘
┌──────────────┼──────────────┬───────────────┐
▼ ▼ ▼ ▼
WebSurfer FileSurfer Coder ComputerTerminal
browser local file writes & executes commands
navigation read/write debugs code and codeWhat is genuinely worth learning in Magentic-One is the Orchestrator's dual-ledger mechanism:
- Task Ledger (outer loop): when the task arrives, write down the known facts, assumptions, and a step-by-step plan; when the inner loop stalls (progress plateaus), return to the outer loop and rewrite the plan instead of grinding on.
- Progress Ledger (inner loop): after every step, evaluate "is the task complete / are we moving forward / who should speak next," and call on the next agent accordingly.
This structure — an explicitly maintained plan and progress, replanning whenever stuck — essentially turns Planning from a prompt trick into an inspectable data structure, and its shadow is visible in many agent systems since, including the various deep research products. On generalist agent benchmarks such as GAIA, AssistantBench, and WebArena, Magentic-One scored comparably to the then-SOTA, making it the representative result of the late-2024 "generalist agent team" route.
In AutoGen 0.4+, it ships as the built-in MagenticOneGroupChat team preset in AgentChat — a few lines of code spin up the same kind of team, still handy for learning and prototyping.
5. AutoGen Studio and .NET Support
AutoGen Studio: a prototyping tool, not a product
AutoGen Studio is the companion no-code GUI: pip install -U autogenstudio, then autogenstudio ui --port 8080, and you can drag components, assemble teams, run tasks, and inspect message traces in the browser. It suits two things: demonstrating multi-agent concepts to non-engineering colleagues, and quickly validating whether a team configuration is sane.
But note two official red lines and one real-world signal:
- The official docs state it is not a production-ready application — it lacks the authentication, security, and other capabilities deployment requires;
- The companion AutoGen Bench is an evaluation suite for running benchmarks, likewise research-oriented;
- The direct consequence of maintenance mode is already visible: as of early 2026, Studio's latest version still depends on
autogen-agentchat<0.6, incompatible with the core library's 0.7.x (GitHub issue #7173 remains unresolved). Dependency drift is itself the most telling symptom of a project entering maintenance mode.
.NET: designed in, but the real answer is MAF
The 0.4 Core layer was designed to support a cross-language .NET/Python runtime in which agents exchange messages across languages. But the .NET-side SDK never left preview, and its feature coverage trails the Python side by a wide margin. As of 2026, Microsoft's official answer for .NET developers is MAF — it treats .NET and Python as first-class (with a Go version besides). .NET teams shouldn't waste evaluation time on AutoGen.
6. Code Examples: 0.4+ Syntax in Practice
The code below is based on autogen-agentchat 0.7.x (install: pip install -U "autogen-agentchat" "autogen-ext[openai]"); all APIs were verified against the official docs and the repo README.
Single agent + tools
python
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.openai import OpenAIChatCompletionClient
def search_docs(keyword: str) -> str:
"""Search the doc library for a keyword and return a summary."""
return f"Search results for {keyword}..."
async def main() -> None:
model_client = OpenAIChatCompletionClient(model="gpt-4.1")
agent = AssistantAgent(
"assistant",
model_client=model_client,
tools=[search_docs], # a plain function is a tool
max_tool_iterations=10, # a single agent may chain up to 10 tool rounds
)
print(await agent.run(task="Look up how termination conditions work"))
await model_client.close()
asyncio.run(main())Note that since 0.6.2, AssistantAgent has its own tool-calling loop via max_tool_iterations, so a single-agent scenario no longer needs a team wrapper.
A writer + critic reflection team (RoundRobin)
python
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient
async def main() -> None:
model_client = OpenAIChatCompletionClient(model="gpt-4.1")
writer = AssistantAgent(
"writer",
model_client=model_client,
system_message="You are the copywriter; revise the draft each round based on the critique.",
)
critic = AssistantAgent(
"critic",
model_client=model_client,
system_message="You are a demanding reviewer; give specific revision requests. Reply APPROVE once the draft is good.",
)
# Semantic termination + a hard-cap backstop, combined with |
termination = TextMentionTermination("APPROVE") | MaxMessageTermination(12)
team = RoundRobinGroupChat([writer, critic], termination_condition=termination)
# Console prints the message stream live to the terminal, including token usage stats
await Console(team.run_stream(task="Write a one-line slogan for a code review tool"))
await model_client.close()
asyncio.run(main())SelectorGroupChat: let the model decide who speaks
python
from autogen_agentchat.teams import SelectorGroupChat
# A three-agent team: after each round, an LLM picks the next speaker
# based on each agent's description
team = SelectorGroupChat(
[planner, coder, reviewer],
model_client=model_client, # picking the speaker itself costs one LLM call
termination_condition=termination,
)Key points: every agent needs a well-written description, since the selection model routes on it; and beyond 4-5 agents, selection quality degrades noticeably — an inherent ceiling of conversational orchestration.
7. The Relationship to Semantic Kernel: Two Frameworks Become Microsoft Agent Framework
For years Microsoft ran two agent frameworks in parallel, divided as "research exploration vs. engineering delivery":
- AutoGen (Microsoft Research): the experimental playground for multi-agent conversation orchestration — fast-moving, aggressive APIs;
- Semantic Kernel (the engineering team): an orchestration SDK for enterprise applications — .NET/Python/Java, emphasizing plugins, telemetry, and stability.
As of the official November 2024 blog post, the line was still "two frameworks in parallel — AutoGen for research, SK for production, with a migration path to come." But the selection confusion and duplicated effort kept growing, and in October 2025 the merger was announced: the two combine into the Microsoft Agent Framework (MAF), with a public preview alongside the .NET ecosystem, RC in February 2026, and 1.0 GA in April 2026. MAF positions itself as the "direct successor" to both: it absorbs AutoGen's clean agent abstractions plus Semantic Kernel's enterprise capabilities (session state management, type safety, middleware, telemetry), and adds graph-style workflows for explicit multi-agent orchestration.
What this means for developers in practice:
- New projects: use MAF directly (
pip install agent-framework); don't start new AutoGen projects; - Existing AutoGen projects: the framework still runs, with bug fixes and security patches maintained by the community, and an official AutoGen → MAF migration guide; the core abstractions (agents, teams, termination) all have counterparts in MAF, so migration is not a rewrite;
- Concept learning: the Team orchestration, termination conditions, and ledger ideas covered on this page all live on in MAF and other frameworks — learning them costs you nothing. For a broader framework comparison, see the framework selection overview, plus the contrasts with LangGraph and CrewAI.
8. Pros, Cons, and Where It Fits
Pros
- The origin of the conversational paradigm, with a simple and direct concept model — agents solve problems by chatting — and the gentlest learning curve among multi-agent frameworks;
- A clean layered design after the 0.4 rewrite: AgentChat for quick prototypes, Core for fine-grained control, each to its own;
- The composable termination-condition design is the most explicit "prevent runaway" mechanism in its class, worth borrowing for every multi-agent system;
- Magentic-One's dual-ledger mechanism is an excellent teaching specimen for planning and multi-agent collaboration;
- A rich ecosystem legacy: oceans of tutorials, paper replications, and community discussion stretching back to the AutoGPT era.
Cons
- It is in maintenance mode: no new features and limited community-maintainer responsiveness — the hardest possible veto;
- Conversational orchestration is structurally expensive: every round broadcasts the full context, and Selector adds another speaker-selection call per round; long tasks consume far more tokens than a single-agent loop (see the cost analysis in Agent Loop);
- Emergent conversation is weakly controllable: with no explicit state graph, termination is the only backstop when agents drift; complex flows are harder to debug than in graph frameworks, and observability has to be built up yourself;
- Studio's version drift against the core library shows the surrounding toolchain is already rusting.
Scenario verdicts
| Scenario | Advice |
|---|---|
| Learning multi-agent concepts, preparing for interviews, replicating papers | Worth learning; the concepts remain current |
| Quickly prototyping a writer-critic pair | Fine to use; AgentChat is the fastest on-ramp |
| New production projects in 2026 (Python) | Choose MAF, LangGraph, or a single agent + tools |
| New production projects in 2026 (.NET) | Go straight to MAF; AutoGen's .NET support never matured |
| Maintaining an existing AutoGen system | Keep it running; schedule the migration to MAF |
The bottom line: AutoGen is worth the time to understand, and not worth investing new projects in. Its historic role — taking multi-agent conversation from a paper to a framework every engineer could actually run — has been inherited by MAF.
References
- microsoft/autogen GitHub repo — the official README, with the maintenance-mode statement, the layered architecture explanation, and 0.4+ quickstart code.
- AutoGen update announcement: merging with Semantic Kernel (Discussion #7066) — the original October 2025 announcement of the merger into Microsoft Agent Framework.
- Microsoft Agent Framework official docs — the MAF overview, explicitly positioned as the direct successor to AutoGen and Semantic Kernel.
- AutoGen AgentChat Teams official tutorial — the authoritative usage of the four Team presets and termination conditions.
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation (arXiv:2308.08155) — the original 2023 framework paper.
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks (arXiv:2411.04468) — the Magentic-One paper, with the dual-ledger mechanism details.
- Magentic-One — Azure AI Foundry Labs — Microsoft's official project page for Magentic-One.
- Microsoft's Agentic Frameworks: AutoGen and Semantic Kernel — the November 2024 blog laying out the two-framework strategy; read it against the post-merger reality.