Skip to content

Glossary

At a glance A collection of 120+ core AI Agent terms organized into seven groups — foundational concepts, reasoning & planning, tools & protocols, memory & retrieval, architecture & orchestration, engineering & operations, and products & ecosystem. Each entry carries a precise one-sentence definition plus links for deeper reading on this site, with anchors for cross-jumping between entries — a pocket dictionary to keep at hand across the whole site.

Glossary ​

This is the dictionary page for the entire site. Whenever you hit an unfamiliar term while reading an article, look it up here first. Each entry gives only a precise one-sentence definition; deeper discussions are reached through the "see" links to the corresponding long-form articles. Entries are grouped into seven thematic sections, with anchors for jumping within and across sections.

How to use this page

  • Entry anchors follow the format #lowercase-hyphenated-english, e.g. [Token](#token). When authoring other pages on this site, you can reference deep links such as /resources/glossary#mcp directly.
  • Definitions reflect industry consensus as of August 2026. Terminology in the Agent field drifts fast — the same word can mean different things a couple of years apart, and key shifts are flagged in the entries.

1. Foundational Concepts ​

  • Agent: A system that perceives its environment, makes decisions autonomously, and invokes tools to take actions toward a goal. In the LLM context, it specifically means a program where the model decides the next step in a loop, rather than following a fixed pipeline. The dividing line with Workflow is "who decides the path". See What is an AI Agent.
  • LLM (Large Language Model): A large neural network trained on massive text corpora whose core task is predicting the next Token — the "engine" of the modern Agent. On its own, though, it only generates text; it is not an Agent.
  • Foundation Model: A general-purpose model pretrained on large-scale data and adaptable to many downstream tasks — GPT, Claude, Gemini, DeepSeek and the like. The emphasis is on the "pretrain once, adapt everywhere" paradigm.
  • Token: The smallest unit of text a model processes, produced by a tokenizer; one token corresponds to roughly 0.75 English words or 0.5–1 Chinese characters. Billing, rate limits, and the context window are all measured in tokens.
  • Context Window: The upper limit on the number of tokens a model can see in a single request, counting both input and output. Everything inside the window is the Agent's "working memory"; anything that doesn't fit must be handled with RAG, Memory, or Compaction. See Context Engineering.
  • Temperature: The parameter that controls sampling randomness — the lower it is, the more deterministic the output. Agent scenarios that write code or call tools usually set it to 0 or near 0; creative writing is where you turn it up. It affects the word-selection distribution, not the model's capability.
  • Prompt: The umbrella term for the input text fed to a model. In Agent systems, a prompt is an engineering artifact dynamically assembled by code, not a hand-written sentence. See Prompt Engineering.
  • System Prompt: The instruction block at the very top of the conversation that defines the model's role, boundaries, and behavioral rules, and takes priority over user messages. System prompts in products like Claude Code run to thousands of tokens and are the main shaper of Agent behavior. This site hosts real system prompts from several products — see the Prompt Archive.
  • In-context Learning: A model's ability to pick up a new task on the fly — without any weight updates — purely from examples or instructions in the prompt. Few-shot (a few examples) and zero-shot (instructions only) are its two flavors, and it is the theoretical basis for every prompt trick.
  • Hallucination: The phenomenon of a model generating plausible-sounding but factually wrong content. The root cause: it is accountable for "does this sound human", not for "is this true". The engineering mitigations are RAG, tool-based verification, and citation tracing — not pleading "please be honest" in the prompt.
  • Reasoning Model: A model specifically trained to produce a long stretch of internal reasoning before answering (OpenAI's o-series, DeepSeek-R1, Claude's extended thinking mode, etc.). It spends more test-time compute to buy accuracy on complex tasks.
  • Inference: The process of generating output with a trained model, as opposed to training. When people say "inference cost", they mean token-generation overhead. Note that "inference" also means logical reasoning — one English word, two senses; tell them apart by context.
  • Fine-tuning: Continuing to train a pretrained model on domain data to change its behavior. The 2026 consensus: if a problem can be solved with prompt + context, don't fine-tune. Reserve fine-tuning for locking in a style, enforcing strict output formats, and distilling small models.
  • RLHF (Reinforcement Learning from Human Feedback): The technique of training a reward model on human preference data and then aligning an LLM through reinforcement learning against it. A later variant, RLAIF, replaces part of the human labeling with AI feedback; Constitutional AI is Anthropic's flagship implementation.
  • Multimodal: A model's ability to handle multiple input/output modalities — text, images, audio, video — in one system. Computer Use relies on the vision modality to read the screen; voice Agents rely on audio for real-time interaction.
  • Structured Output: A feature that forces the model to emit output conforming to a given JSON Schema. It is the foundation of Function Calling and of data exchange between Agents, and far more reliable than "asking nicely for JSON in the prompt".
  • SLM (Small Language Model): A smaller model (on the order of a few billion parameters) that can run locally or at the edge. In Agent systems, SLMs typically handle the cheap front-line steps — routing, classification, guardrail checks — dividing labor with larger models to save cost. See Cost Optimization.
  • Context Engineering: The discipline of designing "what information goes into the context window at each step". It caught fire in mid-2025 through discussions by Shopify CEO Tobi Lütke, Andrej Karpathy, and others, and is seen as prompt engineering in a higher dimension: a prompt is one sentence; context engineering manages the entire information supply system. See Context Engineering.
  • Vibe Coding: A term Karpathy coined in February 2025 for a development style of "barely reading the code, letting AI write it by feel, and cursing another round when it doesn't run". Fine for prototypes and toy projects; a recipe for incidents in production code.
  • AGENTS.md: A file placed at the repo root and written for coding Agents to read — build commands, directory conventions, off-limits areas — essentially a README for new employees. Since 2025 it has been widely supported by OpenAI Codex, Cursor, and Claude Code (whose variant is CLAUDE.md). For how to write one, see Writing AGENTS.md.

2. Reasoning & Planning ​

  • Chain-of-Thought (CoT): The prompting technique of having the model "think step by step" before answering (Wei et al., 2022), which markedly improves performance on math and logic tasks; the starting point of everything later called reasoning. See Core Papers.
  • ReAct: The paradigm proposed by Yao et al. in 2022 (arXiv:2210.03629) that interleaves reasoning (Thought) with action (Action/Observation), letting the model think, look things up, and fix things as it goes. It is the conceptual prototype of the modern Agent Loop and still the most widely used Agent skeleton.
  • Tree of Thoughts (ToT): Expands CoT's single chain into a searchable tree of thought with branching, evaluation, and backtracking (Yao et al., 2023). Beautiful in theory, but inference costs several times a single chain; in practice, most use cases have been taken over by "a reasoning model plus a simple loop".
  • Self-Consistency: Sampling multiple CoT paths for the same question and taking the majority answer by vote. Simple and effective but multiplies the cost; usage has declined since reasoning models took off.
  • Reflection / Self-Critique: The pattern of having the model review its own output, spot errors, and rewrite. Anthropic's building guide lists it among the basic workflows (evaluator-optimizer). It clearly helps writing and code, less so factual errors — models often "fail to reflect on what they don't know".
  • Reflexion: The framework from Shinn et al. (2023) in which, after a failure, the Agent writes the lesson down into memory and reads it back on the next attempt. A landmark of "reinforcement learning with words"; effective for tasks that need repeated trial and error, such as fixing code.
  • Plan-and-Execute: An architecture in which a planner first generates a complete plan and an executor then carries it out step by step (the BabyAGI / Plan-and-Solve lineage). More global awareness than ReAct, but plans go stale easily; in practice people often add "re-plan after every step".
  • Planning: The umbrella term for breaking a goal into ordered subtasks, with implementations ranging from "ask the model to list a plan first in the prompt" to a dedicated planning module. In long-horizon tasks, how often you plan and re-plan decides success or failure. See Planning & Task Decomposition.
  • Task Decomposition: The general technique of splitting big tasks into small ones. CoT (implicit), ToT (explicit search), and Subagents (delegated execution) are all implementations of it.
  • Least-to-Most Prompting: A prompting technique (Zhou et al., 2022) that first has the model decompose the problem into sub-problems ordered from easy to hard, then solves them in sequence. Generalizes better than plain CoT to problems beyond the training distribution.
  • Scratchpad: The technique of letting the model write intermediate computations into its output (Nye et al., 2021), a precursor of CoT. In Agent systems it also refers to the temporary storage an Agent uses to track intermediate state.
  • Reasoning Trace: The intermediate reasoning text a model or Agent produces while solving a problem. Vital for debugging and Evals, but note that traces don't necessarily reflect the model's "real thinking" (the faithfulness problem).
  • Extended Thinking: Anthropic's name for Claude's adjustable reasoning budget, which lets you cap thinking tokens; OpenAI's counterpart is reasoning effort. Both are productized knobs for test-time compute.
  • Test-time Compute: The amount of compute spent at inference time. After o1 in 2024 it became the new scaling axis — instead of making the model bigger, you let it "think longer" to buy performance.
  • Deep Research: A product category in which an Agent autonomously runs multiple rounds of searching, reading, and synthesizing to produce a long, cited research report. OpenAI (February 2025), Gemini, and Perplexity all ship features under this name; at bottom it is "a search-augmented multi-step Agent Loop".
  • Inner Monologue: An early Agent paper's name for the model talking to itself as it reasons (e.g. Inner Monologue, Huang et al. 2022, used in robotics); nowadays generally folded into the term reasoning trace.

3. Tools & Protocols ​

  • Tool Calling: The umbrella term for the mechanism where the model expresses "I want to call this tool" in its output, the host program executes it, and the result is fed back to the model. It is the Agent's only channel to the outside world. See Tools & MCP.
  • Function Calling: The early form and alias of tool calling, in which the model generates a function name and arguments according to a predefined JSON Schema. OpenAI first shipped it in June 2023; other vendors followed and it settled into the umbrella terms tool use / tool calling.
  • Tool Schema: The structured definition of a tool's name, purpose, and parameters in JSON Schema. Writing schemas is a craft — vague descriptions are the number-one cause of misfired tool calls.
  • MCP (Model Context Protocol): The protocol Anthropic open-sourced in November 2024 that standardizes how "LLM apps ↔ external tools/data sources" connect — dubbed "the USB-C of AI". OpenAI and Google adopted it starting March 2025, and in December 2025 it was donated to the AAIF under the Linux Foundation, becoming a vendor-neutral standard. See Tools & MCP.
  • MCP Server / Client / Host: The three roles in MCP. The Server exposes tools, resources, and prompt templates; the Client is the application-side connector, one-to-one with a Server; the Host (Claude Code, Cursor, an IDE, etc.) is the Agent application that initiates connections. Transport supports stdio (local process) and HTTP (remote).
  • A2A (Agent2Agent Protocol): The protocol Google released in April 2025 and donated to the Linux Foundation that June, addressing "how Agents from different vendors discover and cooperate with each other". Complementary rather than competing with MCP — MCP governs Agents calling tools; A2A governs Agents finding Agents. See Multi-Agent Architecture.
  • Agent Card: The JSON "capability card" an Agent publishes about itself under the A2A protocol (conventionally at /.well-known/agent-card.json), declaring what it can do and how to invoke it; the foundation of A2A service discovery.
  • ACP (Agent Communication Protocol): The Agent communication protocol proposed by the IBM/BeeAI camp. After the protocol wars of 2025, the ecosystem converged on A2A, and the ACP lineage's design was folded into the Linux Foundation system. For selection, remember one line: tools are MCP; Agent-to-Agent interop is A2A.
  • Computer Use: The ability of a model to look at screenshots directly and output mouse and keyboard actions to control a real computer. Anthropic shipped the first public API in October 2024, with OpenAI's Operator and others following. Accuracy is still far below human level — treat it as a fallback, not the main channel. See Claude Agent SDK.
  • Browser Agent / Browser Use: The browser-specialized form of Computer Use: parse the DOM or screenshots to drive web pages, commonly for data collection and web automation. The open-source poster child is the browser-use library; on the product side, think of the various Agents that "fill in forms and book tickets for you".
  • Code Interpreter: Giving the model a Sandbox in which it can execute code, so it "writes code to solve the problem" instead of doing mental arithmetic. A qualitative leap for math and data analysis; now a standard tool in ChatGPT, Claude, and peers.
  • GUI Agent: The umbrella term for Agents that operate graphical interfaces; Computer Use and Browser Agents are its subtypes. The core difficulty is visual grounding — mapping "the OK button" to pixel coordinates.
  • Agent Skills: A capability packaging format introduced by Anthropic in October 2025 and released as an open standard in December (spec at agentskills.io): a folder containing a SKILL.md (metadata + instructions) plus optional scripts and reference files, loaded by the Agent on demand. Think of it as a "reusable pack of expert experience".
  • Progressive Disclosure: The token-saving design principle shared by Skills and MCP: at startup, load only the name and description (a few dozen tokens); read the full instructions when a task matches; pull in reference files on demand during execution — three-tier loading that keeps the resident context minimal.
  • Tool Registry: The catalog in which an Agent system registers, discovers, and manages available tools. The MCP ecosystem launched an official Registry in 2025 to solve server discovery and trust; enterprises also commonly build internal registries as the single point of control for permissions and auditing.

4. Memory & Retrieval ​

  • RAG (Retrieval-Augmented Generation): The architecture that first retrieves relevant passages from a knowledge base, then feeds them to the model together with the question to generate an answer. The main weapon against hallucination and knowledge cutoff, and still the first choice for enterprise deployments in 2026. See RAG.
  • Agentic RAG: The form of RAG in which the Agent autonomously decides "whether to search, what to search, how many rounds, and whether it's enough" — as opposed to naive single-shot retrieval. Higher quality but slower and more expensive; suited to complex questions.
  • Embedding: The technique of mapping text (or images) into high-dimensional vectors such that semantically similar items land close together; the mathematical foundation of semantic retrieval. When choosing an embedding model, check the MTEB leaderboard and language coverage — not just the dimension count.
  • Vector Database: A database that stores embeddings and supports approximate nearest neighbor (ANN) search — think Pinecone, Milvus, Qdrant, pgvector. At small-to-medium scale, pgvector is often enough; there is no need to jump straight to a dedicated vector store.
  • Semantic Search: Search based on vector similarity rather than keyword matching. Strong on paraphrases, weak on exact terms and codes — which is why production systems almost always mix it with BM25.
  • Chunking: The strategy of splitting long documents into retrieval-sized pieces. Chunks that are too big dilute relevance; too small lose context. The mainstream approach cuts at semantic boundaries (headings, paragraphs) and keeps overlap; advanced options include parent-child chunks and late chunking.
  • HyDE (Hypothetical Document Embeddings): The technique (Gao et al., 2022) of first having the model fabricate a "hypothetical answer" and then using its vector to retrieve real documents. Effective when questions and answers are phrased very differently, at the cost of one extra generation of latency.
  • Reranking: The stage in which a stronger cross-encoder model re-ranks the retrieved candidates. It usually lifts retrieval quality by a full notch — one of the highest-value RAG optimizations.
  • Hybrid Search / BM25: BM25 is the classic keyword-statistics retrieval algorithm; hybrid search fuses BM25's exact matching with vector-based semantic matching, then refines the results with rerank. The default recipe for production RAG in 2026.
  • Knowledge Graph: A graph structure that stores knowledge as entity–relation–entity triples. In Agent scenarios it serves multi-hop reasoning and explainable retrieval; costly to build, and best suited to relation-dense domains (healthcare, legal, finance).
  • GraphRAG: The method Microsoft proposed in 2024: use an LLM to extract a knowledge graph from a corpus, then answer global questions from community summaries. Strong at "what is this corpus about overall" questions; painful incremental updates are its main weakness.
  • Long-term Memory: User preferences, facts, and experience persisted across sessions, as opposed to short-term memory within a single session (the conversation history). Implementations are usually a three-part stack: write-decision logic, vector/structured storage, and recall-time injection. See Memory Systems.
  • Memory Types: A three-way scheme borrowed from cognitive science — episodic memory (what happened), semantic memory (knowledge and facts), procedural memory (how to do things). In Agent systems these map to conversation archives, a fact store, and Skills/SOPs.
  • Context Compaction: The technique of summarizing early history into a digest when a session nears the context window limit to free up space; the /compact command in Claude Code and peers is the manual version. Summaries are necessarily lossy — persist key decisions before compaction eats them.
  • Context Rot / Context Poisoning: The phenomenon where the longer the context and the more irrelevant or wrong information mixed in, the worse the model performs. Multiple studies in 2025 confirmed that "a big window doesn't mean you can use it all well". The remedies are stuffing less, clearing often, and compacting in time — not mindlessly piling up to 1M tokens.
  • Memory Tool: A design that turns memory operations themselves into tools (save_memory / search_memory) the model invokes autonomously. More token-efficient than "inject the full memory every turn", and closer to how humans behave — write it down when it comes up.

5. Architecture & Orchestration ​

  • Workflow: The LLM orchestration style in which steps and control flow are hard-coded by the developer, as opposed to an Agent (where the model decides the path dynamically). Anthropic's advice: if a workflow can do the job, don't reach for an Agent — workflows win across the board on determinism, cost, and debuggability. See Concept Distinctions.
  • Agent Loop: The minimal Agent skeleton — the cycle of "model thinks → calls tool → observes result → thinks again", until the task completes or a termination condition fires; every fancy architecture is a variation on this loop. See Agent Loop.
  • Orchestration: The umbrella term for organizing and scheduling flows of multiple LLM calls, tools, and Agents. The framework wars (LangGraph vs CrewAI vs hand-rolled) are, at bottom, a fight over orchestration abstractions. See Framework Overview.
  • Multi-Agent System: An architecture in which multiple Agents, each with its own specialty, divide the work. The gains are context isolation and specialization; the costs are debugging pain and token burn several times over (Anthropic measured roughly 4–15× a single Agent). See Multi-Agent Architecture.
  • Supervisor: A multi-Agent topology in which one central Agent assigns tasks and aggregates results while the others just work and don't talk to each other. The most used and most reliable multi-Agent structure; both the OpenAI Agents SDK and LangGraph support it out of the box.
  • Handoff: The mechanism by which one Agent transfers control of a conversation, together with its context, to another Agent. The core primitive of the OpenAI Agents SDK; "transfer to a human / transfer to a specialist" in customer service is the canonical handoff.
  • Subagent: A temporary Agent spawned by the main Agent with its own context window, bringing back only its conclusions when done. Claude Code's Task tool is exactly this design — isolating context and keeping the main thread clean is its biggest value.
  • Router: The pattern of spending one cheap call to classify intent, then dispatching the request to different branches (models, prompts, or Agents). The simplest orchestration optimization, and the one you should do first.
  • Orchestrator-Workers: The topology (Anthropic's naming) in which a central LLM dynamically decomposes a task, dispatches pieces to parallel workers, and synthesizes the results. Differs from Supervisor in that the subtasks are generated at runtime.
  • Evaluator-Optimizer: The workflow in which one LLM generates, another LLM scores against criteria and gives feedback, and the two loop until the bar is met. The engineered version of the Reflection pattern; fits tasks with clear acceptance criteria.
  • Parallelization: The pattern of running multiple LLM calls at the same time and merging the results, in two flavors: sectioning (parallel subtasks) and voting (parallel attempts). What you buy is not just speed but accuracy too (majority vote).
  • HITL (Human-in-the-Loop): The design of pausing execution at critical junctures (deleting data, sending email, making payments) to await human approval or correction. Not a crutch for weak capability but the default architecture in high-risk scenarios. See Human-in-the-Loop.
  • State: The data structure carried across steps while an Agent runs (message history, intermediate artifacts, counters). LangGraph treats state as a first-class citizen: all nodes read and write the same typed state.
  • Checkpoint: The mechanism that persists an Agent's intermediate state, enabling resume after interruption and time travel. LangGraph's checkpointer is the reference implementation, and the technical precondition that makes HITL "pause and wait for a human" possible.
  • Graph-based Agent: Modeling an Agent's flow as a directed graph of nodes and edges — nodes are computation, edges are transition conditions. More explicit and persistable than a pure code loop, at the price of an abstraction tax; using it for a simple Agent is over-engineering.
  • Agent-as-Tool: The pattern of wrapping a complete Agent as an ordinary tool for another Agent to call. Nearly synonymous with Subagent; different frameworks use different names (LangGraph calls it tool-calling subgraph, OpenAI calls it agents as tools).
  • Swarm: A decentralized multi-Agent topology in which Agents communicate peer-to-peer with no central scheduler. OpenAI's experimental teaching framework Swarm (2024) popularized the word (its ideas later merged into the Agents SDK), but production systems almost always use Supervisor; swarm remains mostly a research concept.

6. Engineering & Operations ​

  • Eval: The engineering practice of measuring Agent quality with structured test sets and scoring criteria. Iterating on an Agent without evals is driving with your eyes closed — the sharpest line separating engineering teams from toy projects in 2026. See Evaluation Systems and Evals in Practice.
  • Benchmark: A public, standardized evaluation set (e.g. SWE-bench, GAIA). Good for cross-model comparison, but no substitute for a private eval tailored to your own business — a high leaderboard score doesn't mean it performs on your tasks.
  • LLM-as-a-Judge: The evaluation method of having another LLM score outputs as a judge. Scalable, but with its own biases (it prefers long answers and its own output style), so calibrate the judge itself periodically against human labels.
  • Trace: The complete record of one Agent run — every step's input, output, tool calls, token spend, and timing; the primary crime scene for debugging Agents. See Observability.
  • Observability: The practice of collecting traces, metrics, and logs from an Agent system to understand its behavior. Representative tools: LangSmith, Langfuse, Braintrust, Arize Phoenix; since 2025, OpenTelemetry's GenAI semantic conventions have been the de facto standard at the collection layer.
  • Prompt Injection: An attack in which malicious instructions are hidden in external content — web pages, emails, documents — to hijack any Agent that reads it. The number-one threat in Agent security, and there is no perfect fix today; the only defense is defense in depth: least privilege, isolation, and human confirmation. See Security & Alignment.
  • Jailbreak: A conversational attack that coaxes a model into bypassing its own safety limits. Differs from prompt injection in that the target is the model itself rather than the external content it reads.
  • Guardrails: Programmatic checks on both sides of the model's input and output (sensitive words, PII filtering, topic restrictions, format validation). The point: guardrails run outside the model, deterministic and auditable — not yet another "please don't make mistakes" prompt.
  • Sandbox: An isolated environment (containers, microVMs like Firecracker, gVisor) for executing Agent-generated code or commands. Giving an Agent shell access without a sandbox is handing production machines to an intern who can be prompt-injected.
  • Red Teaming: The practice of actively playing attacker and systematically probing an Agent's security boundaries. Mandatory before launch, and it must cover tool-abuse paths, not just chat content.
  • Least Privilege: Granting each Agent tool only the minimum permissions needed to finish its task (read-only over read-write, scoped directories over full disk). The last wall that keeps an injection from turning into a catastrophe.
  • Approval Flow: The mechanism requiring a human or a policy engine to approve high-risk actions before they execute. In the same family as HITL; engineering implementations typically lean on checkpoints to pause and resume.
  • Data Exfiltration: The attack in which an injected Agent ships out sensitive data it has read through tool calls (firing requests, writing to public repos). The key defense is restricting outbound channels, not counting on the model to "behave".
  • Golden Dataset: A carefully human-labeled test set kept as a regression baseline. It needn't be large (dozens to a few hundred cases), but it must cover the core business paths and keep growing with real-world bad cases.
  • Rate Limit: The caps model APIs place on request frequency and token throughput (RPM/TPM). A single Agent task can fire dozens of calls, so Agents hit limits far more easily than chat apps; build queues and backoff into the architecture.
  • TTFT (Time to First Token): The time from sending a request to receiving the first token — the core metric for streaming experience. For Agents, what matters more is end-to-end task duration: TTFT plus generation time, accumulated over multiple rounds.
  • Prompt Caching: The mechanism by which vendors cache the KV of repeated prefixes (system prompts, long documents) to cut price and latency (supported by Anthropic, OpenAI, and Gemini). The engineering rule: stable content first, variable content last; the hit rate determines the size of your bill. See Cost Optimization.
  • Retry & Fallback: The recovery strategy when a model call fails — retry transient errors with exponential backoff, and on persistent failure degrade to a backup model or a simpler prompt. A multi-model fallback chain is standard equipment for production Agents.
  • Determinism / Reproducibility: The property of producing the same output for the same input. Even at temperature=0, LLMs don't guarantee token-for-token identical output (a consequence of floating point and batching), so evals and regression tests should assert on "stable in distribution" rather than "exactly identical".

7. Products & Ecosystem ​

  • Claude Code: Anthropic's terminal-native coding Agent (GA in May 2025), known for its minimal toolset + long system prompt + Subagents; the benchmark of "less is more" Agent design, and its runtime has since been extracted as the Claude Agent SDK. See Claude Code Anatomy.
  • Claude Agent SDK: The development framework Anthropic extracted from the Claude Code kernel (formerly Claude Code SDK), suited to building Agents whose objects of operation are files and the shell. See Claude Agent SDK.
  • Cursor: Anysphere's AI-native IDE, which evolved from "an editor with Copilot" into a development environment centered on a built-in Agent mode; the representative of the IDE-route coding Agent. See Cursor Anatomy.
  • GitHub Copilot: GitHub's AI coding assistant, which started with inline completions and since 2025 has gone all-in on Agent mode (async coding agents, assignable issues). Its edge is deep integration with the GitHub workflow (PRs, Actions, Code Review).
  • Devin: Cognition's "AI software engineer", sold on fully autonomous completion of development tasks. Its 2024 launch demo sparked enormous controversy and pushed the industry to draw a harder line between demos and deliverable capability. See Devin Retrospective.
  • Manus: The general-purpose Agent product released in March 2025 by the Butterfly Effect (Monica) team, which broke out with its form of "a complete computer environment running in a cloud VM" and set off China's general-purpose Agent boom. See Manus Anatomy.
  • OpenAI Codex: OpenAI's coding Agent product line (CLI + cloud sandbox tasks, launched 2025); note that it is a different thing from the early code model of the same name from 2021.
  • LangChain: The LLM application framework founded by Harrison Chase in 2022 — the largest ecosystem, and also the most criticized for its layers of abstraction. Version 1.0 (LTS) shipped in October 2025, with the core converging on create_agent and middleware and the underlying runtime ceding to LangGraph.
  • LangGraph: The LangChain team's graph-structured Agent orchestration framework: flows modeled as state + nodes + edges, with checkpointing and HITL built in. With 1.0 in October 2025 it became one of the production-grade default choices. See LangGraph in Practice.
  • OpenAI Agents SDK: OpenAI's lightweight Agent framework (the successor to the experimental Swarm project), whose core primitives number just four: Agent, Handoff, Guardrails, and Tracing. The philosophy: few primitives, close to the model API. See OpenAI Agents SDK.
  • Responses API: OpenAI's Agent-oriented API launched in March 2025, building tool calling, reasoning items, and conversation state management into the interface; it is gradually displacing Chat Completions as the default entry point for new applications.
  • CrewAI: The multi-Agent framework built on the "role-playing team" metaphor, with each Agent carrying a role/goal/backstory. Quick to pick up and concept-friendly; a good fit for business-process prototypes. See CrewAI.
  • AutoGen / AG2: Microsoft Research's open-source multi-Agent conversation framework, centered on "conversation between Agents is programming". The community forked AG2 in late 2024, and in 2025 Microsoft folded its capabilities into later AutoGen versions and Agent Framework — mind the version genealogy when choosing. See AutoGen.
  • LlamaIndex: The framework strongest at data connectors and indexing, with the most complete RAG pipeline capabilities; it later grew Agent and workflow modules as well. A common pick for knowledge-intensive Agent projects.
  • Dify: The open-source low-code platform for LLM applications (LangGenius), with visual orchestration of workflow + RAG + Agent; friendly to self-hosting and widely adopted for private deployments in China. See Dify.
  • Coze: ByteDance's Bot/Agent building platform; its moat is the plugin ecosystem and distribution channels (Doubao, Feishu, WeChat). Well suited for operations and non-engineering roles to ship Bots fast. See Coze Anatomy.
  • n8n: The open-source workflow automation tool; since 2024 it has shipped built-in AI Agent nodes (with tool calling and memory), making it a popular pick for hybrid "traditional automation + LLM" orchestration.
  • OpenHands: The open-source AI software developer platform (formerly OpenDevin), providing a complete Agent runtime, sandbox, and evaluation infrastructure — the standard experimental playground for coding Agent research. See OpenHands.
  • SWE-bench: The benchmark released by a Princeton team in 2023 that evaluates Agents' bug-fixing ability with real GitHub issues. The official family includes the original (2294 tasks, 12 Python repos), Lite (300), Verified (500, human-screened), Multilingual (300, 9 languages), and Multimodal (517). Verified has been driven close to saturation by top models, and in 2026 the field has moved to the harder SWE-bench Pro and Terminal-Bench. See SWE-agent Case Study.
  • Terminal-Bench: The benchmark (Stanford University and Laude Institute et al., 2025) that evaluates Agents on system-level tasks in a real terminal — provisioning environments, running services, writing scripts. Closer than SWE-bench to "full-stack ops" capability.
  • GAIA: The general-assistant benchmark (2023) released jointly by Meta, HuggingFace, AutoGPT and others, testing web browsing, multimodality, and tool-chain coordination in combined tasks. Tasks come in three difficulty tiers; it was once the touchstone for general Agents, and top systems now score close to its ceiling.
  • AAIF (Agentic AI Foundation): The dedicated fund established by the Linux Foundation in December 2025, co-founded by Anthropic, Block, and OpenAI with support from Google, Microsoft, AWS and others, hosting Agent infrastructure projects such as MCP and goose; it marks the arrival of vendor-neutral governance for the Agent protocol layer.

Terms expire

Some of the pre-2025 buzzwords on this page — AutoGPT-style "fully autonomous Agents", GPT-3-era "prompt magic" — have already become historical relics, while terms that only emerged in 2024–2025, such as MCP, Skills, and context engineering, are now interview staples. The right way to treat academic jargon is to understand the problem it solves, not to recite definitions — the engineering trade-offs behind a concept are the long-term asset. See A Brief History.

References ​