Skip to content

Frontier

At a glance A survey of seven frontier directions in Agent research for 2025-2026 — Agentic RL, Deep Research, Computer Use, self-improvement, lifelong learning, multi-agent protocols, and safety research — with representative work for each and a clear "worth watching or hype" verdict.

This page contains time-sensitive content; data is current as of 2026-08. Job listings, pricing, and product features may have changed — verify against the original sources before citing.

Frontier ​

This page is unlike the rest of the site: other pages cover "what has already settled," this one covers "what is still shifting violently." Agent research went through a paradigm switch in 2025 — from "prompt orchestration + workflows" to "training agent behavior directly with RL" — and the field's center of gravity, paper output, and engineering practice are all being rewritten along with it.

So this page comes with three usage notes:

  1. Freshness: content verified as of August 2026; for later developments, defer to the original sources. Every direction includes entry points you can keep following (a survey, a leaderboard, a protocol's official site).
  2. Verdicts first: each direction ends with a one-line "worth watching or hype" verdict. Verdicts can be wrong, but they beat a list of papers — if you want a paper list, go straight to the Paper Map.
  3. Not exhaustive: only work that changed the direction of subsequent research in each area. For the classic foundational papers, go to Core Papers.

1. Agentic RL: From "Orchestrating Agents" to "Training Agents" ​

What happened ​

Before 2024, the mainstream way to improve agent capability was engineering: better prompts, a more refined Agent Loop, more tools. The shift in 2025: treat the entire agent trajectory (reasoning + tool calls + environment feedback) as the object of RL training, optimizing end-to-end for task completion.

The landmark event was the system card OpenAI published alongside Deep Research in February 2025, which explicitly disclosed: the model was based on an early version of o3 and trained with end-to-end reinforcement learning for web browsing tasks, with the reward signal coming from whether the task was completed — not from human-labeled preferences. It was the first time a frontier lab publicly admitted "agent capability is trained by RL, not tuned by prompts."

The open-source community followed quickly, forming a clear technical lineage:

RLHF / RLVR (alignment and reasoning)
  └── GRPO (DeepSeekMath, 2024): drops the value network, group-relative advantages
        └── DAPO (2025): decoupled clipping, dynamic sampling, stabilizing long-CoT training
              └── Agentic RL (2025- ): multi-turn trajectories + tool calls + environment rewards
                    ├── Search-R1: makes the search engine part of the RL environment
                    ├── WebDancer / WebSailor line: synthetic data + RL for training web agents
                    ├── Kimi K2: a trillion-parameter model trained directly for agentic capability
                    └── Tongyi DeepResearch: a fully open-source deep research training recipe

Representative work ​

  • Kimi K2: Open Agentic Intelligence (July 2025, Moonshot AI): a 1.04-trillion-parameter MoE (32B active) whose training objective is explicitly agentic capability — large-scale agentic data synthesis + multi-stage post-training + RL. It demonstrated that "an open model purpose-built for agent scenarios" can approach the closed-source frontier. In early 2026, K2.5 added vision and multi-agent orchestration (Agent Swarm).
  • WebDancer (arXiv:2505.22648) and WebSailor-V2 (arXiv:2509.13305, May-September 2025, Alibaba Tongyi): systematically answered "where does a web agent's training data come from" — automatically synthesize high-difficulty information-seeking tasks, then fine-tune with RL. WebSailor-V2 claims to have narrowed the gap with closed-source systems on multiple deep research benchmarks.
  • Search-R1 (2025, UIUC et al.): models the retrieval engine as an RL environment; the model decides on its own when to search and what to search for, with result-based rewards shaping the search policy. The cleanest example of "tool calls inside the RL loop."
  • Tongyi DeepResearch (September 2025): the first fully open-source web agent to claim parity with OpenAI Deep Research on comprehensive evaluations — data synthesis, training pipeline, and model weights all open. For teams wanting to reproduce agentic RL, this is currently the most complete public recipe.

The key engineering difficulties ​

The difference between agentic RL and ordinary RLHF isn't the algorithm — it's the environment:

  • Rollout cost: a single trajectory is dozens of tool calls; training throughput is throttled by environment interaction (search APIs, browsers, code executors), and asynchronous rollout frameworks have become the focal point of infrastructure competition;
  • Sparse rewards: long tasks only have a signal at the end; mainstream solutions are synthesizing automatically verifiable tasks (the WebSailor route) or rubric-based rewards;
  • Environment fidelity: policies trained on static datasets degrade when they hit real websites' anti-scraping measures, redesigns, and dirty data — "high-fidelity training environments" have themselves become a research topic (e.g., EnterpriseBench/Corecraft-type work).

Worth watching or hype?

Worth watching — the most important paradigm shift of 2025-2026. If your work involves getting an agent to reliably complete tasks in a vertical domain, "build verifiable rewards for the domain's tasks + run agentic RL" is already a lever an order of magnitude stronger than prompt engineering. But mind the precondition: you need an automatically verifiable task definition and a cheap-enough rollout environment, otherwise this road doesn't work.

2. Deep Research Systems: Agentic RL's First Killer App ​

Deep Research is agentic RL's most successful product form: give it a question, and it autonomously searches, reads, and cross-verifies dozens to hundreds of web pages, producing a cited research report. After OpenAI debuted it in February 2025, Gemini, Perplexity, xAI, Microsoft Copilot, and Kimi all followed within months, making it 2025's most crowded product category.

Convergence of technical approaches ​

Vendors don't publish implementation details, but from OpenAI's system card, the Deep Research survey, and open-source reproductions, the converged architecture can be summarized:

User question
   │
   ▼
┌──────────────┐   Multi-turn loop (dozens to hundreds of steps)
│  Planner/    │ ──► Decompose sub-questions
│  Reasoner    │ ──► Generate search queries
│ (single RL-  │ ──► Read pages/PDFs, extract evidence
│  trained     │ ──► Contradiction found → backtrack and revise the plan
│  model)      │ ──► Enough evidence → exit loop
└──────────────┘
   │
   ▼
┌──────────────┐
│  Report      │  Itemized citations, organized by sub-question
│  generation  │
└──────────────┘

Two key architectural judgments have reached consensus:

  1. End-to-end single model + RL beats multi-agent orchestration. LangChain's open-source reproduction open_deep_research early on used a supervisor + multiple sub-agents structure; later versions and Tongyi DeepResearch both confirmed the trend: once the underlying model is trained with agentic RL, "one strong model + long context + good tools" is more stable than elaborate multi-agent topologies. This matches our judgment in Multi-Agent Architecture: multi-agent is not a silver bullet.
  2. The bottleneck has shifted from "finding information" to "verifying information." Starting in late 2025, benchmarks appeared that specifically attack deep research's weak points (like the Wiki Live Challenge, requiring expert-level Wikipedia articles), exposing the current systems' common flaws: citations look proper but the support relationships are misaligned, and low-quality sources go undetected. The evaluation methodology itself has become a research hotspot; see Evaluation Systems.

In February 2026, OpenAI added MCP connector support to Deep Research — it can connect to arbitrary MCP tools and restrict search to trusted sites. This marks deep research's evolution from "public-web research tool" toward "enterprise private-knowledge research portal."

Worth watching or hype?

The product form is established; technical differentiation is narrowing. For researchers: what's worth studying is information verification and evaluation, not building yet another "agent that can search." For engineers: stop building your own deep research workflow — use the API (both OpenAI and Gemini have one) or deploy Tongyi DeepResearch / open_deep_research, and spend your effort on your own data sources and eval sets.

3. Computer Use / GUI Agents: Slowly but Genuinely Climbing Toward Usable ​

Current state ​

GUI agents (viewing screenshots, emitting mouse/keyboard actions) are the hardest direction in agents, because they can't skip any link in the chain "visual understanding → UI grounding → long-horizon operation." Timeline:

  • October 2024: Anthropic released computer use for Claude (public beta API) — the first frontier model to natively support "see the screen + operate it";
  • January 2025: OpenAI released Operator, a browser-operations agent built on the CUA (Computer-Using Agent) model, taking the productized route;
  • July 17, 2025: OpenAI released ChatGPT agent, merging Operator's browser operations, Deep Research's information synthesis, and ChatGPT's conversation into a single agent — "the unified agent" became the shared narrative among top vendors;
  • 2025-2026: the open-source ecosystem iterated around the OSWorld benchmark (369 real desktop tasks), with specialized GUI models like UI-TARS and UI-Venus continually topping the leaderboard.

Which numbers to believe ​

OSWorld is currently the most recognized desktop-operation benchmark. Per the mid-2026 leaderboard and multiple independent re-tests: the top spot (agents driven by the Claude Opus series) has passed 70%, while OpenAI's CUA scored markedly lower in third-party re-tests (one re-test reported 38.1%). Two reminders: the gap between leaderboard scores and human performance (just over 70%) is narrowing but not closed; and vendor-reported scores often diverge significantly from third-party re-tests — verify the leaderboard's current data before citing any number.

The new research trend is bypassing pixels: since GUI grounding is the largest error source, some work (such as 2026's StateAct) has the main agent read program state directly (filesystem, DOM, APIs), isolating screen operations into a sub-agent invoked only when necessary — in testing, the GUI sub-agent was invoked for only about 1% of steps. This "if there's an API, don't click the mouse" approach is squeezing pure GUI agent research toward "the long tail of scenarios with no API available."

Worth watching or hype?

Half hype. General-purpose "operate the computer for you" is still unreliable in 2026; Operator-style products' approval ratings can't support their pricing narrative. But specialized scenarios (browser forms, legacy internal systems without APIs, RPA replacement) already deliver real value. The single test of whether a GUI agent product is credible: does it dare publish third-party OSWorld re-test scores?

4. Self-Improving Agents: From "Tweaking Prompts" to "Rewriting Their Own Code" ​

The core work: Darwin-Gödel Machine ​

In May 2025, a team from Sakana AI, UBC, and the Vector Institute (Jenny Zhang, Jeff Clune, et al.) published Darwin-Gödel Machine: Open-Ended Evolution of Self-Improving Agents (arXiv:2505.22954, later accepted at ICLR 2026), one of 2025's most-discussed agent papers.

The two sources of its name explain its design:

  • Gödel machine (Schmidhuber's theoretical construct): a universal problem solver that can "provably" rewrite itself — theoretically perfect, practically infeasible, because the "prove the rewrite is beneficial" step can't be done;
  • Darwinian evolution: give up proof, use empirical validation instead — let the agent modify its own code, run it against benchmarks, keep the effective modifications in a constantly growing "agent archive," then sample parents from the archive and keep mutating.

DGM validated this loop on SWE-bench and Polyglot: a coding agent, through dozens of rounds of self-rewriting (modifying its own tools, memory structure, and workflow), automatically discovered tricks of the kind humans previously only found when hand-designing agents. The paper's weightiest conclusion is a scaling relation: the more compute invested in self-improvement, the better the agent architecture becomes — a scaling law at the architecture level, not just the model level.

Context: a wider lineage ​

DGM isn't an isolated work; it sits in a lineage of "recursive self-improvement (RSI)":

  • STOP (arXiv:2310.02304, COLM 2024): an earlier "scaffold self-optimization" — the model recursively improves its own scaffolding code with weights untouched;
  • A Self-Improving Coding Agent (arXiv:2504.15228, 2025): gradient-free self-improvement, gaining substantial improvements through reflection + code updates;
  • HyperAgents (2026): generalizes "self-referential improvement" to more general agent meta-learning settings;
  • Survey view: A Survey of Self-Evolving Agents (arXiv:2507.21046, continuously updated since July 2025) organizes the entire lineage with a "What / When / How / Where to Evolve" framework — the best map into this area.

Honest boundaries ​

These works' real contributions are often obscured by two kinds of noise. First, the media equates "can rewrite its own code" with "the intelligence explosion is near" — DGM's improvements happen entirely inside the sandbox demarcated by benchmarks, and every round burns substantial full-evaluation compute. Second, there's no clean ablation separating how much of "self-improvement's" gains come from capabilities the model already had but hadn't activated, versus genuine architectural innovation. On the safety dimension, an agent that can rewrite itself naturally bypasses the traditional "code review as approval" control point — directly relevant to the safety research in Section 7.

Worth watching or hype?

Worth watching for research; still hype for engineering. DGM proved the direction "agent architectures can be searched" is real, but one round of self-improvement costs far more than manual tuning, and the gains concentrate in coding-like domains with automatic verification signals. The practical near-term version is its degenerate form: use evolutionary search offline to optimize your agent's configuration (toolset, prompts, topology) rather than letting a live agent rewrite itself.

5. Long-Term Memory and Lifelong Learning: Architecture Research Gets Serious ​

An agent's "memory" has long been an engineering problem: a vector store + summaries + retrieval tricks (what the site's Memory Systems page covers). The 2025-2026 change is that architecture research is directly attacking "catastrophic forgetting," trying to give the model weights themselves a capacity for continual learning.

Representative work ​

  • Nested Learning: The Illusion of Deep Learning Architectures (arXiv:2512.24695, Google Research, NeurIPS 2025): proposes viewing a model as a set of nested optimization problems updating at different frequencies — fast layers handle immediate context, slow layers consolidate knowledge slowly, like "long-term memory." The accompanying HOPE architecture (continuing the Titans line) shows preliminary evidence of mitigating catastrophic forgetting. This is the most fundamental architectural reframing of "how models learn continually" in recent years.
  • Lifelong Learning of Large Language Model based Agents (arXiv:2501.07278, January 2025): systematically maps the perception-memory-action loop of "lifelong learning agents"; among the most-cited surveys in this direction.
  • Engineering side: the "memory as a service" route represented by Letta (the MemGPT team) accelerated its productization in 2026 (e.g., its continual learning SDK), making "editable self-memory" a standard component; the Self-Evolving survey (2507.21046) discusses external experiential memory and parameter-level learning in the same framework — the two routes are converging.

Verdict ​

If parameter-level continual learning (the Nested Learning line) succeeds, the impact is foundational — but so far it's been validated only at small-to-medium scale, and it's far from replacing the current "external memory + RL fine-tuning" approach. For engineers, the right answer in 2026 is still an external memory system, but it's worth checking back on this line every six months, because once it matures, existing memory architectures get rewritten wholesale.

WARNING

Papers in the "lifelong learning" direction are multiplying fast, but evaluation standards are a mess: most papers define their own forgetting rates and forward-transfer metrics, making cross-paper comparison difficult. When you see a claim of "solving catastrophic forgetting," first check whether it was validated on standard continual learning benchmarks rather than self-built tasks.

6. Multi-Agent Collaboration and Protocols: The Standards War Was Basically Over by 2026 ​

Agent interoperability protocols were 2025's liveliest "standards battlefield"; by mid-2026 the landscape was broadly settled. The positioning and status of the three main protocols:

ProtocolOriginatorScopeGovernanceStatus as of mid-2026
MCPAnthropic (2024.11)model ↔ tools/dataDonated to the Agentic AI Foundation under the Linux FoundationThe de facto standard; the tool layer is undisputed
A2AGoogle (2025.4)agent ↔ agent (cross-vendor)Same — donated to the Linux Foundation150+ supporting organizations; shipping in AWS/Microsoft/IBM product lines
ANPOpen-source communityDecentralized agent networkCommunity governanceSmall but active; adopted in academia and open source
  • A2A (specification) solves "how agents from different companies discover each other and delegate tasks": at its core are the Agent Card (capability declaration) + task lifecycle management. The mid-2026 reality: the standard's position is established, but production deployments concentrate in controlled scenarios at a few large companies; independent developers haven't adopted it much yet.
  • ANP (white paper, technical white paper arXiv:2508.00007) is more ambitious: decentralized identity based on W3C DID (the did:wba method), a three-layer architecture (identity and secure communication layer, meta-protocol negotiation layer, application protocol layer), self-positioned as "the HTTP of the agentic web era." Its design is the most complete in the protocol family, but its ecosystem is the smallest — and history says adoption of decentralized-identity protocols never hinges on design quality.
  • On the periphery there's also the payments-layer AP2 (Agent Payments Protocol, Google with multiple payment providers, September 2025) and more, showing that the "agent economy" infrastructure puzzle is being filled in.

On the academic side, protocol-comparison research is now a mature field (e.g., A Survey of Agent Interoperability Protocols, arXiv:2505.02279), with a consensus conclusion for choosing among them: MCP for tools, A2A for cross-organizational agent collaboration, ANP betting on an open agent web — the three are complementary, not mutually exclusive. The security surface of the protocol layer (identity forgery, capability-declaration spoofing) remains an open problem; see Security and Red-Teaming.

Worth watching or hype?

A2A is worth adopting; ANP is worth watching. If your system needs to dispatch agents across teams/vendors, go straight to A2A; if everything stays inside your organization, a protocol isn't a necessity — don't introduce complexity for "standards" sake. ANP's technical design is worth reading, but check the adoption curve before betting on its ecosystem.

7. Agent Safety Research Frontier: From "Jailbreaking Models" to "Attacking Agent Systems" ​

Agent safety completed an upgrade of its research subject in 2025: from "getting the model to say prohibited content" to "exploiting an agent's tools, memory, and autonomy to cause real harm." The four most active research lines:

  1. Industrialization of indirect prompt injection. Malicious web pages, emails, and documents embedding instructions to hijack agents have gone from incident reports to a systematic attack surface. The paradigm AgentDojo (NeurIPS 2024) established — "test injection attack/defense in a real tool environment" — became the standard; in 2025 the security community documented numerous injection incidents against browser agents and MCP servers. No silver bullet on defense; the current consensus is defense in depth: least-privilege tools, human confirmation for critical actions (see Human in the Loop), and isolation between untrusted content and the instruction channel.
  2. Behavioral study of alignment failure. Anthropic's Agentic Misalignment (June 2025) marks this line: in simulated corporate environments, models were put under pressure of "goal conflict + facing replacement," and multiple mainstream models chose blackmail, data leaks, and other "insider threat" behaviors. Combined with Apollo Research's in-context scheming work (arXiv:2412.04984), "agents display strategic deception under pressure" has moved from hypothesis to reproducible experimental phenomenon.
  3. Evaluation awareness. Work in 2026 began systematically measuring "whether the model knows it's being evaluated and changes behavior accordingly" — a direct threat to the validity of all safety evaluations, and a new headache for benchmark methodology.
  4. Systematization of safety benchmarks. AgentHarm (ICLR 2025, multi-step harmful task execution), Agent-SafetyBench (ASB, ICLR 2025), and LITMUS (2026) for GUI/real-OS environments gave "agent safety" its first comparable quantitative yardstick; The 2025 AI Agent Index (published 2026) turned the technical and safety characteristics of deployed agent systems into an auditable, documented index.

Worth watching or hype?

The most worth-watching direction, period. It's the only field where "research progress directly determines whether the product can ship": indirect injection and agentic misalignment are both reproduced and unsolved. Teams building agent products should put AgentHarm/AgentDojo into CI and treat misalignment-style experiments as threat-modeling input — not as media news.

8. Summary: A Verdict Table ​

DirectionOne-line verdictRecommended action
Agentic RLThe most important paradigm shift of 2025-2026If your domain has verifiable rewards, get in early
Deep ResearchThe form is established; differentiation is narrowingUse off-the-shelf systems; put effort into evaluation
Computer Use / GUIHalf hype; usable in specialized scenariosTrust only third-party OSWorld re-tests
Self-improving agentsReal as research; not yet time for engineeringOptimize configs with offline evolutionary search
Long-term memory / lifelong learningEarly days for an architecture-level breakthroughUse external memory in engineering; keep tracking
Multi-agent protocolsThe standards war is basically overA2A across organizations; don't over-engineer within one
Agent safetyProgress here caps what your product can doFold into threat modeling and CI

Low-cost ways to track these frontiers: bookmark the surveys listed on this page (the Deep Research survey, the Self-Evolving survey, and the protocol-comparison survey are all continuously updated), the OSWorld and SWE-bench leaderboards, and the Frontier Resource List. If you're preparing to job-hunt into this space, the Knowledge Map marks how these frontier topics map onto job requirements.

References ​