Skip to content

AutoGen

At a glance Microsoft Research's conversational multi-agent framework — the Core/AgentChat/Extensions three-layer architecture after the 0.4 rewrite, Team orchestration modes, and the Magentic-One system — plus the real 2026 selection advice now that AutoGen has been folded into Microsoft Agent Framework.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

AutoGen ​

1. What It Is — and the 2026 Reality You Must Know First ​

AutoGen is the multi-agent framework open-sourced by Microsoft Research, with the paper posted and code released in August 2023 (arXiv:2308.08155). It founded the "conversational multi-agent" school: multiple agents with distinct roles collaborate on a task through natural-language conversation, rather than being driven by a predefined state graph. Its core abstractions influenced nearly every multi-agent framework that followed, and the "two agents talk their way to a result" paradigm went viral starting here.

But discussing AutoGen in August 2026 requires putting one overwhelming fact on the table first:

AutoGen has entered maintenance mode

In October 2025, Microsoft announced that AutoGen and Semantic Kernel would merge into the unified Microsoft Agent Framework (MAF); MAF shipped its 1.0 GA in April 2026. The AutoGen repo's README now states plainly: no new features will be accepted, the project is community-maintained, new projects should use MAF directly, and existing projects should migrate per the official migration guide.

That doesn't make this page worthless, for three reasons:

  1. A massive amount of existing code and tutorials. Multi-agent tutorials, paper replications, and interview questions from 2023-2025 are full of AutoGen APIs — you will need to read them sooner or later.
  2. The concepts are still universal. Team, termination conditions, speaker selection — MAF and other frameworks inherited these concepts wholesale, so learning AutoGen means learning the intellectual source of the whole school.
  3. Interviewers love it. Multi-agent orchestration is a high-frequency topic in agent job descriptions, and "what's the difference between AutoGen's GroupChat and LangGraph's graph orchestration" is a classic question.

The one-line positioning: study it as the living fossil and idea library of multi-agent architecture — don't pick it as a production framework for new projects in 2026.

2. Version History: The 0.2-to-0.4 Ground-Up Rewrite ​

AutoGen's history splits cleanly into two eras:

The 0.2 era: the classic conversation API (2023-2024) ​

The old API took ConversableAgent as its base class; typical code looked like this:

python
# old 0.2 API (deprecated; shown only so you can recognize old tutorials)
from autogen import AssistantAgent, UserProxyAgent

assistant = AssistantAgent("assistant", llm_config={...})
user_proxy = UserProxyAgent("user_proxy", code_execution_config={...})
user_proxy.initiate_chat(assistant, message="Draw me a stock price chart")

AssistantAgent proposes ideas, UserProxyAgent stands in for the human and executes code, and initiate_chat steps on the gas to start the conversation; multi-party scenarios use GroupChat + GroupChatManager. This API was extremely easy to grasp and single-handedly popularized the multi-agent concept, but the engineering problems piled up: synchronous blocking, messy typing, hidden global state, painful debugging, and no fine-grained control over message flow.

How to tell old tutorials from new at a glance

from autogen import ... or initiate_chat(...) means the old 0.2 API, deprecated since 2025; the new API imports entirely from the three packages autogen_agentchat, autogen_core, and autogen_ext. Most tutorials still online are the old kind — copying them will trip you up.

The 0.4 era: the full rewrite (released January 2025) ​

In January 2025, Microsoft shipped AutoGen 0.4 — officially described as a rewrite from scratch. The core changes:

Dimension0.20.4+
Execution modelSynchronous, blocking conversation loopFully async, event-driven (async/await)
PackagingSingle package pyautogenSplit into autogen-core / autogen-agentchat / autogen-ext
ArchitectureFlat; agents call each other directlyLayered: Core runtime → AgentChat high-level API → Extensions
ObservabilityEssentially noneBuilt-in tracing, message streaming, OpenTelemetry
Cross-languagePython onlyThe Core layer is designed for a cross-language .NET/Python runtime
TypingWeakFully type-annotated

The layered architecture can be understood like this:

┌───────────────────────────────────────────────────────────────┐
│  App layer      AutoGen Studio (no-code GUI) · Magentic-One   │
├───────────────────────────────────────────────────────────────┤
│  AgentChat     AssistantAgent · Teams · termination conditions│
│                 ← the only layer most people ever need        │
├───────────────────────────────────────────────────────────────┤
│  Core           event-driven actor runtime · messaging        │
│                 ← drop down here only for fine-grained control│
└───────────────────────────────────────────────────────────────┘
        Extensions cut across every layer: model clients / code execution / MCP tools
  • Core: an event-driven actor runtime where agents communicate only through async messages; supports local and distributed deployment. Flexible but verbose — the layer for framework-level customization.
  • AgentChat: the "opinionated" high-level API on top of Core — preset agents and preset teams, the first choice for quick prototypes, and the closest to the 0.2 mental model.
  • Extensions: pluggable implementations — model clients (OpenAI, Azure OpenAI, etc.), the Docker code executor, MCP workbenches, and more.

Iteration continued after 0.4; the last stable minor line before maintenance mode was 0.7.x (autogen-agentchat 0.7.5, released September 2025), requiring Python ≥ 3.10. No feature-bearing major version followed.

3. Core Concepts: Agents, Teams, and Termination ​

The AgentChat layer in 0.4+ has only three core concepts; grasp them and you've grasped the framework.

AssistantAgent and the built-in agents ​

AssistantAgent is the workhorse: it wraps an LLM and can carry tools, a system message, and memory. Other built-ins include UserProxyAgent (human in the loop — waits for human input each round), CodeExecutorAgent (executes code and returns the results), and MultimodalWebSurfer (browser operation). Tools are ordinary Python functions whose schemas are generated automatically from type annotations and docstrings; you can also attach MCP servers directly (see Tools & MCP).

Teams: four preset orchestration modes ​

A Team is a group of agents plus a set of rules governing "who speaks next and when to stop." AgentChat ships four presets:

TeamSpeaker-selection mechanismBest for
RoundRobinGroupChatFixed order, taking turnsDeterministic flows like two-agent reflection (writer + critic)
SelectorGroupChatAn LLM picks the next speaker each roundFree-form discussion with many roles and no fixed flow
MagenticOneGroupChatAn Orchestrator assigns work dynamically via a ledgerOpen-ended web/file/code tasks (next section)
SwarmAgents hand off proactively via HandoffMessageSupport-style "transfer" flows

RoundRobinGroupChat is cheap and predictable — the default starting point. SelectorGroupChat is flexible but adds one extra LLM call per round just to choose the speaker, so cost and latency climb noticeably as rounds accumulate.

Single agent first, team later

The official AutoGen docs themselves urge this: teams need more steering scaffolding, so optimize a single agent's tools and instructions first, and move to a team only once you've proven the single agent isn't enough. This matches the conclusion of the Multi-Agent Architecture page — for most tasks multi-agent is over-engineering, and a writer-critic pair already captures most of the real benefit.

Termination conditions: the gate that keeps conversations from burning your budget ​

The biggest engineering risk in conversational orchestration is "talking forever" — agents being polite to each other, correcting each other, the token bill climbing exponentially. AutoGen makes termination conditions a first-class citizen, combinable with bitwise operators:

python
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination

# Stop when the critic says APPROVE; but force a stop after 20 messages no matter what (the backstop)
termination = TextMentionTermination("APPROVE") | MaxMessageTermination(20)

Also common: TimeoutTermination (by time), TokenUsageTermination (by token usage), and ExternalTermination (stopped by an external signal). The iron rule of engineering practice: a semantic condition (such as a keyword) plus a hard cap (message count / tokens) always appear as a pair. This is also the most basic line of defense for cost and budget control in multi-agent scenarios.

4. Magentic-One: the Blueprint for a Generalist Multi-Agent System ​

In November 2024, Microsoft Research's AI Frontiers lab released Magentic-One (paper arXiv:2411.04468), built on AutoGen and positioned as a "generalist" multi-agent system: instead of tuning prompts for specific tasks, a lead Orchestrator coordinates four specialist agents to complete open-ended web, file, and code tasks.

              ┌─────────────────┐
              │   Orchestrator  │  Outer loop: Task Ledger (known facts + the plan)
              │   (lead agent)  │  Inner loop: Progress Ledger (progress + next step)
              └────────┬────────┘
        ┌──────────────┼──────────────┬───────────────┐
        ▼              ▼              ▼               ▼
   WebSurfer      FileSurfer        Coder       ComputerTerminal
   browser        local file        writes &    executes commands
   navigation     read/write        debugs code and code

What is genuinely worth learning in Magentic-One is the Orchestrator's dual-ledger mechanism:

  • Task Ledger (outer loop): when the task arrives, write down the known facts, assumptions, and a step-by-step plan; when the inner loop stalls (progress plateaus), return to the outer loop and rewrite the plan instead of grinding on.
  • Progress Ledger (inner loop): after every step, evaluate "is the task complete / are we moving forward / who should speak next," and call on the next agent accordingly.

This structure — an explicitly maintained plan and progress, replanning whenever stuck — essentially turns Planning from a prompt trick into an inspectable data structure, and its shadow is visible in many agent systems since, including the various deep research products. On generalist agent benchmarks such as GAIA, AssistantBench, and WebArena, Magentic-One scored comparably to the then-SOTA, making it the representative result of the late-2024 "generalist agent team" route.

In AutoGen 0.4+, it ships as the built-in MagenticOneGroupChat team preset in AgentChat — a few lines of code spin up the same kind of team, still handy for learning and prototyping.

5. AutoGen Studio and .NET Support ​

AutoGen Studio: a prototyping tool, not a product ​

AutoGen Studio is the companion no-code GUI: pip install -U autogenstudio, then autogenstudio ui --port 8080, and you can drag components, assemble teams, run tasks, and inspect message traces in the browser. It suits two things: demonstrating multi-agent concepts to non-engineering colleagues, and quickly validating whether a team configuration is sane.

But note two official red lines and one real-world signal:

  • The official docs state it is not a production-ready application — it lacks the authentication, security, and other capabilities deployment requires;
  • The companion AutoGen Bench is an evaluation suite for running benchmarks, likewise research-oriented;
  • The direct consequence of maintenance mode is already visible: as of early 2026, Studio's latest version still depends on autogen-agentchat<0.6, incompatible with the core library's 0.7.x (GitHub issue #7173 remains unresolved). Dependency drift is itself the most telling symptom of a project entering maintenance mode.

.NET: designed in, but the real answer is MAF ​

The 0.4 Core layer was designed to support a cross-language .NET/Python runtime in which agents exchange messages across languages. But the .NET-side SDK never left preview, and its feature coverage trails the Python side by a wide margin. As of 2026, Microsoft's official answer for .NET developers is MAF — it treats .NET and Python as first-class (with a Go version besides). .NET teams shouldn't waste evaluation time on AutoGen.

6. Code Examples: 0.4+ Syntax in Practice ​

The code below is based on autogen-agentchat 0.7.x (install: pip install -U "autogen-agentchat" "autogen-ext[openai]"); all APIs were verified against the official docs and the repo README.

Single agent + tools ​

python
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.openai import OpenAIChatCompletionClient

def search_docs(keyword: str) -> str:
    """Search the doc library for a keyword and return a summary."""
    return f"Search results for {keyword}..."

async def main() -> None:
    model_client = OpenAIChatCompletionClient(model="gpt-4.1")
    agent = AssistantAgent(
        "assistant",
        model_client=model_client,
        tools=[search_docs],          # a plain function is a tool
        max_tool_iterations=10,       # a single agent may chain up to 10 tool rounds
    )
    print(await agent.run(task="Look up how termination conditions work"))
    await model_client.close()

asyncio.run(main())

Note that since 0.6.2, AssistantAgent has its own tool-calling loop via max_tool_iterations, so a single-agent scenario no longer needs a team wrapper.

A writer + critic reflection team (RoundRobin) ​

python
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient

async def main() -> None:
    model_client = OpenAIChatCompletionClient(model="gpt-4.1")

    writer = AssistantAgent(
        "writer",
        model_client=model_client,
        system_message="You are the copywriter; revise the draft each round based on the critique.",
    )
    critic = AssistantAgent(
        "critic",
        model_client=model_client,
        system_message="You are a demanding reviewer; give specific revision requests. Reply APPROVE once the draft is good.",
    )

    # Semantic termination + a hard-cap backstop, combined with |
    termination = TextMentionTermination("APPROVE") | MaxMessageTermination(12)
    team = RoundRobinGroupChat([writer, critic], termination_condition=termination)

    # Console prints the message stream live to the terminal, including token usage stats
    await Console(team.run_stream(task="Write a one-line slogan for a code review tool"))
    await model_client.close()

asyncio.run(main())

SelectorGroupChat: let the model decide who speaks ​

python
from autogen_agentchat.teams import SelectorGroupChat

# A three-agent team: after each round, an LLM picks the next speaker
# based on each agent's description
team = SelectorGroupChat(
    [planner, coder, reviewer],
    model_client=model_client,          # picking the speaker itself costs one LLM call
    termination_condition=termination,
)

Key points: every agent needs a well-written description, since the selection model routes on it; and beyond 4-5 agents, selection quality degrades noticeably — an inherent ceiling of conversational orchestration.

7. The Relationship to Semantic Kernel: Two Frameworks Become Microsoft Agent Framework ​

For years Microsoft ran two agent frameworks in parallel, divided as "research exploration vs. engineering delivery":

  • AutoGen (Microsoft Research): the experimental playground for multi-agent conversation orchestration — fast-moving, aggressive APIs;
  • Semantic Kernel (the engineering team): an orchestration SDK for enterprise applications — .NET/Python/Java, emphasizing plugins, telemetry, and stability.

As of the official November 2024 blog post, the line was still "two frameworks in parallel — AutoGen for research, SK for production, with a migration path to come." But the selection confusion and duplicated effort kept growing, and in October 2025 the merger was announced: the two combine into the Microsoft Agent Framework (MAF), with a public preview alongside the .NET ecosystem, RC in February 2026, and 1.0 GA in April 2026. MAF positions itself as the "direct successor" to both: it absorbs AutoGen's clean agent abstractions plus Semantic Kernel's enterprise capabilities (session state management, type safety, middleware, telemetry), and adds graph-style workflows for explicit multi-agent orchestration.

What this means for developers in practice:

  • New projects: use MAF directly (pip install agent-framework); don't start new AutoGen projects;
  • Existing AutoGen projects: the framework still runs, with bug fixes and security patches maintained by the community, and an official AutoGen → MAF migration guide; the core abstractions (agents, teams, termination) all have counterparts in MAF, so migration is not a rewrite;
  • Concept learning: the Team orchestration, termination conditions, and ledger ideas covered on this page all live on in MAF and other frameworks — learning them costs you nothing. For a broader framework comparison, see the framework selection overview, plus the contrasts with LangGraph and CrewAI.

8. Pros, Cons, and Where It Fits ​

Pros ​

  • The origin of the conversational paradigm, with a simple and direct concept model — agents solve problems by chatting — and the gentlest learning curve among multi-agent frameworks;
  • A clean layered design after the 0.4 rewrite: AgentChat for quick prototypes, Core for fine-grained control, each to its own;
  • The composable termination-condition design is the most explicit "prevent runaway" mechanism in its class, worth borrowing for every multi-agent system;
  • Magentic-One's dual-ledger mechanism is an excellent teaching specimen for planning and multi-agent collaboration;
  • A rich ecosystem legacy: oceans of tutorials, paper replications, and community discussion stretching back to the AutoGPT era.

Cons ​

  • It is in maintenance mode: no new features and limited community-maintainer responsiveness — the hardest possible veto;
  • Conversational orchestration is structurally expensive: every round broadcasts the full context, and Selector adds another speaker-selection call per round; long tasks consume far more tokens than a single-agent loop (see the cost analysis in Agent Loop);
  • Emergent conversation is weakly controllable: with no explicit state graph, termination is the only backstop when agents drift; complex flows are harder to debug than in graph frameworks, and observability has to be built up yourself;
  • Studio's version drift against the core library shows the surrounding toolchain is already rusting.

Scenario verdicts ​

ScenarioAdvice
Learning multi-agent concepts, preparing for interviews, replicating papersWorth learning; the concepts remain current
Quickly prototyping a writer-critic pairFine to use; AgentChat is the fastest on-ramp
New production projects in 2026 (Python)Choose MAF, LangGraph, or a single agent + tools
New production projects in 2026 (.NET)Go straight to MAF; AutoGen's .NET support never matured
Maintaining an existing AutoGen systemKeep it running; schedule the migration to MAF

The bottom line: AutoGen is worth the time to understand, and not worth investing new projects in. Its historic role — taking multi-agent conversation from a paper to a framework every engineer could actually run — has been inherited by MAF.

References ​