Skip to content

Agents and Dialogue Systems

On this page Dialogue policy learning, tool-use and web-navigation agents, embodied agents — RL fine-tuning that treats the LLM as the policy; the latest progress from ReAct to RL-trained agents (e.g., AgentLM, RFT); the environment is an API.

Agents and Dialogue Systems ​

In a nutshell: this page explains agent RL that treats the LLM as the policy — from dialogue-policy learning in early dialogue systems to tool use in the ReAct paradigm, web navigation, and embodied agents. The core questions: what does the environment interface look like, where does the reward come from, and why does the "task success rate" objective change the shape of RLHF.

1. The Agent Paradigm as an MDP: Environment = API / Browser / World ​

1.1 Defining the Agent ​

After 2022, "agent" became the central concept in LLM applications: an LLM wrapped as a looping decision-making unit that repeatedly executes "read state → decide action → take action → observe the outcome":

text
┌──────────────┐  action (call API / click page / issue tool call)   ┌───────────────┐
│  LLM agent   │ ──────────────────────────────────────────────────► │  Environment  │
│  (policy π_θ)│ ◄───────────────────────────────────────────────── │  API/browser/ │
└──────────────┘   observation (tool return / page state / error)    │  real world   │
                                                                     └───────────────┘

In MDP terms (see the Markov decision process page):

text
State s_t    : dialogue history + tool returns + environment observations (page DOM, file system, robot sensors)
Action a_t   : generate a passage of text that may contain a tool call (e.g., a JSON-formatted function call)
Reward r_t   : task success, tool-call validity, user feedback (sparse!)
Termination  : task completed, step budget exhausted, or the agent gives up

1.2 Three Key Differences (vs. Traditional RLHF) ​

DimensionRLHF (chat alignment)Agent RL
RewardHuman preference (RM score)Task completion (judged by rules/executor)
EnvironmentThe model itselfExternal APIs / browsers / the physical world
SparsityOne terminal scoreEqually sparse, but "success" can be checked automatically
StateAutoregressive textMultimodal, long context, partially observable

This shift to "reward = task success" is the key to understanding agent RL: when the reward can be verified programmatically, there's no need to train a reward model — and that leads straight to RLVR (see the frontier advances page).

2. Dialogue Policy Learning: RL's Long History in Dialogue Systems ​

2.1 The "Policy" of Task-Oriented Dialogue Systems ​

Before LLMs, task-oriented dialogue systems (flight booking, weather lookups) used a pipeline architecture: NLU → dialogue policy → NLG. The dialogue policy decides "what the system should say or do next" — naturally an RL problem:

text
State  : dialogue state (user intent, slot-filling status)
Action : system acts (ask for a slot, confirm, query the database, give the final reply)
Reward : +1 for task completion, a small cost per dialogue turn, a heavy penalty for wrong confirmations

Zhao & Eskenazi (2016) used DQN to learn dialogue policies, showing that deep RL can learn better "ask–answer–confirm" strategies than hand-written rules. Earlier work (the statistical dialogue systems of the 2010s) had also formalized the problem as a POMDP.

2.2 Historical Lesson: Why Dialogue-Policy RL Never Took Off ​

  • Simulators are hard to build: a user simulator is a fiction, and policies learned against it degrade on real human conversations;
  • Sparse, noisy rewards: whether the task succeeded is hard to determine automatically;
  • Data scarcity: real dialogue data is limited.

This history hands us a general law: the ceiling of RL = environment fidelity × reward quality. The reason LLM agents revived dialogue RL is precisely that they swapped the "environment" from "user simulators" to "executable APIs and browsers" — rewards can finally be verified automatically.

3. The ReAct Paradigm and RL Fine-Tuning of LLM Agents ​

3.1 ReAct: Making the LLM Explicitly "Think → Act → Observe" ​

ReAct (Reasoning + Acting), proposed by Yao et al. (2022), is now the standard skeleton for agents: the LLM alternately outputs "Thought → Action → Observation." This format separates "reasoning" from "acting" explicitly, greatly improving debuggability and success rates.

text
Thought: I need to check the weather in Beijing
Action: weather_api(city="Beijing")
Observation: {"temp": 28, "condition": "sunny"}
Thought: It's sunny, so I can recommend going out
Final: It's sunny in Beijing today, 28°C — great weather for being outdoors.

ReAct itself is a prompting technique; it doesn't train the model. What was truly revolutionary is that it provided a unified action format — which paved the way for "training agents with RL."

3.2 RL Fine-Tuning of Agents: AgentLM and Friends ​

When ReAct prompting falls short (long tasks, complex tools, frequent errors), research turns to RL fine-tuning of the LLM to make its tool calls more reliable. Representative work:

  • AgentLM (Zeng et al., 2023): build a large-scale tool-call dataset, do SFT first, then further alignment, improving multi-step tool-call success rates;
  • RFT (rejection sampling fine-tuning): sample a large number of answers from the model, filter the "successful" subset with rewards or rules, then SFT on that subset — essentially "offline reward maximization," simpler and more stable than PPO;
  • ToolLLM (Qin et al., 2023): expand the tool-call tree with depth-first search, then train tool use with RFT.

3.3 Reward Sources: A Ladder of Executability ​

Reward design for tool-call RL follows a clear ladder:

Reward levelWhat it coversCharacteristics
Process rewardIs the tool-call syntax valid; was a sensible tool calledDense; computable automatically
Outcome rewardWas the tool return used correctly; is the final answer rightMedium density
Task rewardDid the whole multi-step task succeedSparse; most truthful

The key engineering judgment: process rewards (e.g., "this API call was well-formed") let RL quickly learn "how to call," but may produce idle behaviors like "calling the right API to no effect"; task rewards are the most truthful but the sparsest. Industry practice leans toward a mix: process rewards to guide, task rewards to gate.

3.4 The ReAct Family: From Reflection to Multi-Round Improvement ​

ReAct opened a family of training-free methods that "let the LLM judge itself," complementing RL:

MethodMechanismRelationship to RL
Reflexion (2023)After a failure, the agent writes "why it failed" into verbal memory so as to avoid it next round"Verbal RL": improvement via language instead of gradients
Self-Refine (2023)Generate → self-critique → revise, iterating until satisfiedTraining-free "internalization of the reward signal"
Tree search (ToT/DFS)Search multiple paths over the reasoning treeIsomorphic to AlphaGo's search × learning
Agent fine-tuning with RLActually update model weights with verifiable rewardsThe "gradient version" of the above two

Practical judgment: for small tasks, Reflexion/Self-Refine are enough (zero cost); turn to RL when "reflection prompts no longer work and the weights must change." This follows the "start simple" principle on the RL design principles page.

3.5 A Complete Agent MDP Example: Tool Calling ​

Let's write "an LLM calls tools to complete a task" as a trainable MDP and see each component's role:

text
Task: look up a stock price for the user and generate an up/down analysis

State s_t    : dialogue history (user instruction + generated thoughts/actions/observations) + tool catalog
Action a_t   : generate a passage of text (containing a function call such as stock_price(symbol="AAPL"))
Transition   : the tool returns JSON -> appended to the observations
Reward r_t   : (terminal) does the final answer contain the correct price ± analysis-quality score
Termination  : a Final answer is given / the step budget is exceeded

Training data: real user instructions x tool-call trajectories x terminal success/failure

Three agent-specific difficulties this example exposes:

  1. States are "semi-structured": dialogue history mixes natural language with JSON, beyond the fixed-length state assumption of ordinary RL;
  2. Actions are "generative": the action space is text space, not an enumerated action set — the "action probabilities" in RL must live on the token distribution;
  3. Rewards are extremely sparse: only terminal success or failure, with no score for intermediate steps — calling for process rewards or curriculum decomposition.

Together, these three difficulties explain why agent RL is harder than both RLHF and classic RL, and why "get prompting working first → then RFT → finally PPO/GRPO" is the industry's recognized escalation path.

4. RL for Web Navigation and Code Generation ​

4.1 Web Navigation ​

Task: given an instruction ("book me a flight from Beijing to Shanghai"), the agent must complete operations on real or simulated web pages (clicking, typing, navigating).

  • WebArena / WebArena-Lite (Zhou et al., 2023): self-hosted environments running real web applications, evaluating agents' multi-step web operation skills; they have become the de facto standard in this direction;
  • Training recipe: first collect expert trajectories offline for SFT, then RL fine-tune on "whether the final task succeeded";
  • Difficulty: web states are partially observable (page content is extremely long; scrolling or clicking is needed to reveal it), so the agent must learn "which operations reveal which information."

4.2 Code Generation ​

When "writing code" is the task, the reward can be verified directly by an executor:

text
State  : problem statement + code generated so far + compilation/test results
Action : the next stretch of code (a token sequence)
Reward : unit-test pass rate / passing hidden tests / code execution results

Tasks with "verifiable rewards" like this make RL unusually clean — no RM training needed; test cases serve as the judge. This is the same thread behind the explosive growth of "RL for reasoning" (math, code) in 2024–2025; see the RLVR section on the frontier advances page for details.

4.3 Training Details for Code-Generation RL ​

RL fine-tuning for code tasks has four engineering points:

PointPracticePitfall
Test-set splitStrict separation of train/validation/test casesTest-case leakage = score inflation
Reward granularityUnit-test pass rate vs. +1 only if all tests passGrade partial passes, but don't let RL learn to "cherry-pick the easy ones"
Number of samples8–64 answers sampled per problemToo few samples and nothing is learned; too many and costs explode
Compilation environmentExecute code in a sandbox against malicious inputNo sandbox = disaster (see the safety note below)

WARNING

Code-execution safety is a hard boundary: an RL-trained agent executes its own generated code inside the environment. If that code can reach the file system or the network, a malicious prompt mixed into the training data could teach the agent destructive behaviors. In industry, execution must be isolated with Docker sandboxes, timeouts, and resource limits. This is cognate with the "safety guardrails are non-negotiable" principle on the Robotics Control and Sim2Real page.

TIP

Code and math are the "ideal laboratory" for agent RL: rewards are automatically verifiable, environments can be repeated indefinitely, and failures can be precisely attributed. If you're new to agent RL, start with "code completion / unit tests" or "math problem solving" — far easier than open-ended web navigation. For how "verifiable rewards" change RLHF, revisit the RLVR section of LLM Alignment: RLHF in Practice.

5. Embodied Agents and VLA: Agents and Robotics Converge ​

5.1 VLA: Vision-Language-Action Models ​

Since 2023, the embodied agent direction has merged the LLM's "general language understanding" with the robot's "action output" into a single model (VLA, Vision-Language-Action):

text
Input   : camera images + language instructions
Output  : robot actions (joint position deltas / end-effector poses)
Training: massive robot teleoperation data + internet data -> SFT, then RL fine-tuning
  • RT-2 / RT-2-X (Google DeepMind, 2023): improve VLA generalization with web-scale image-text data;
  • OpenVLA (2024): an open-source VLA that the community can fine-tune;
  • RL fine-tuning makes VLAs more effective at "learning from failure" — but with real robot data scarce, it still happens mostly in simulation.

5.2 The Connection to Sim2Real ​

RL for embodied agents and Robotics Control and Sim2Real are two ends of the same thing:

  • Robotics side: low-level continuous control (joint torques) is learned with PPO/SAC in simulation;
  • Agent side: high-level task planning ("pick up the cup first, then pour the water") is learned with LLM + RL;
  • The shared challenge: environment fidelity and reward quality. Whether "learned pouring" in simulation survives a real kitchen depends on the simulator's physics and scene coverage.

6. Environment Engineering: API Wrapping and Reward Sources ​

6.1 "The Environment Is an API" ​

Almost all of the engineering weight in agent RL sits on the environment side:

ComponentEngineering taskCommon pitfall
Action spaceWrap environment operations into a unified tool schema (name/params/returns)Inconsistent schemas cause hallucinated calls
Observation returnsStructure tool results (JSON/tables) and cap their lengthOverlong returns blow the context
Reward judgingA task-success judge (rules/executor/LLM judge)The judge itself is unreliable
Reset and samplingEnvironments must be resettable and sampleable in parallelStateless environments; Docker isolation needed
Error handlingTimeouts, retries, graceful degradation on tool failureFrequent timeouts during training wreck efficiency

DANGER

Environment bugs are the number-one invisible killer of agent RL: a tool occasionally returning corrupted JSON, a page element occasionally failing to click — these teach the agent quirks like "retry over and over" or "just give up," while the training curve looks perfectly healthy. This matches "environment bugs silently poison learning" on the common pitfalls page exactly, except agent-environment bugs are even harder to debug (they involve multiple services, networks, and permissions). Pre-launch environment health monitoring and failure-injection testing are mandatory.

6.2 Three Routes to Reward Sources ​

  1. Rules/executors (test cases, API response validation) — the most reliable; use automatic verification whenever you can;
  2. LLM judges (a strong LLM answers "did this response complete the task?") — flexible but gameable; calibration needed;
  3. Human feedback (the traditional RLHF route) — the most truthful but the most expensive; usually reserved for the final stage.

6.3 Engineering Details of Action-Space Design ​

In agent RL, the "action space" is a collection of tool calls, and its design quality decides whether training succeeds:

Design pointGood practiceBad practice
Tool granularityOne tool per API, with an explicit JSON Schema for the parametersCramming all capabilities into one giant tool
Number of toolsExpose 10–50 at a time, dynamically loaded per taskHanding over hundreds at once; the agent freezes with indecision
Return formatStructured (JSON/tables) + truncation + error codesLong unstructured natural-language returns
Failure semanticsTool failures return readable errors instead of throwingSilent failures (the agent thinks it succeeded)
Context managementSummarize historical observations; keep key stateContext balloons without bound; the model loses focus

The easiest pitfall to fall into is "dirty success/failure signals in tool returns" — the agent learns the bad habit of "calling the tool without looking at the result." If you want a trainable environment, first make tool returns stable, structured, and clearly success/failure-distinguishing. This aligns with the "the environment is the product" thesis on the Anatomy of an RL System page.

7. The Evaluation Problem: Beyond Success Rate ​

7.1 Why "Task Success Rate" Alone Isn't Enough ​

Evaluating agents only by "did it succeed in the end" hides a host of problems:

DimensionWhy it mattersCommon distortion
Step efficiencySucceeding in 20 steps vs. 5 steps is a different qualityLong detours still count as success
Failure modesIs it "can't call tools" or "misread the task"?Success rate alone can't attribute blame
GeneralizationDoes it still work on a different task distribution or site layout?In-distribution score inflation
SafetyDid it take dangerous actions (deleting files, placing orders)?The highest-success agent may be the most dangerous
RobustnessHow many retries until success?A one-off success may be luck

Reporting protocol: task success rate + average steps + failure-mode distribution + multi-seed variance — none can be omitted. This is the protocol from the evaluation and benchmarks page, deployed in the agent setting.

7.2 The Generalization Dilemma ​

The most painful bottleneck in agent RL right now is generalization: an agent trained on WebArena falls apart on a different website. The reason is that the distribution of web pages and APIs is far too broad for any training distribution to cover. Industry routes forward:

  • Larger-scale sampling from diverse environments;
  • Few-shot adaptation (show the agent a few examples of the new API before letting it work);
  • Write "generalization" into the training objective itself (the domain-randomization idea — see Robotics Control and Sim2Real).

7.3 A Map of Agent Evaluation Benchmarks ​

You can't evaluate agents only on your own environment; know the mainstream public benchmarks:

BenchmarkWhat it testsCharacteristics
WebArena (2023)Multi-step web tasks (booking, shopping, admin)Self-hosted real applications; reproducible tasks
SWE-bench (2023)Fixing code from real GitHub issuesDirectly tests "engineering ability"; hard and real
GAIA (2023)General-assistant multi-step tasksTests reasoning + tools + multimodality combined
ToolBench / API-BankTool-calling abilityTests API selection and call correctness

Usage: a paper's "success rate" for an agent must state which benchmark and which protocol were used; when comparing across benchmarks, first check whether the task difficulty is comparable. This also echoes the "benchmark map + reporting standards" methodology on the evaluation and benchmarks page.

8. Frontiers and Criticisms ​

8.1 Frontier Directions ​

  • Scaling up RLVR: DeepSeek-R1 proved that pure RL (no SFT cold start) can elicit reasoning; agent versions are following;
  • Self-improvement loops: agents using feedback to improve their own prompts or policies (Reflexion et al.), naturally complementary to RL's "online sampling";
  • Multi-agent agents: multiple LLM agents collaborating on complex tasks — which circles back to the non-stationarity problems of multi-agent RL.

8.2 Criticisms and Cold Water ​

  • Opaque evaluation: many agent papers report success rates on self-built environments, making cross-paper comparison poor;
  • "RL training" vs. "prompt engineering": some papers' "RL fine-tuning" is actually just RFT (rejection-sampling SFT), not true policy gradients;
  • Cost: every sample in agent RL requires multi-step environment interaction, and the token cost scales with the failure rate — training agents with RL costs an order of magnitude more than alignment.

8.3 Cost Accounting: What Agent RL Actually Costs ​

Do the math first, then decide whether to adopt agent RL:

text
Cost of one training sample = tokens per task interaction x number of attempts x unit price

Example (web-navigation task, 7B model, internal pricing $0.3 per million tokens):
  5,000 tokens per task on average (input + output) x 32 samples = 160k tokens per task
  A 10k-task training set -> 1.6B tokens ≈ $480 (sampling only)
  Add the training itself (a few epochs each of SFT/PPO) plus failure retries -> multiply by another 2–4x
StageCost sourceControl lever
Data samplingMultiple rollouts per taskCap the sample count; validate on a small task set first
Environment interactionWeb/API calls, simulatorsCache reusable trajectories; parallelize
TrainingPPO/GRPO multi-model forward passesGet it working with a smaller model before scaling up
Failure retriesRework caused by environment instabilityEnvironment health monitoring; failure injection

Engineering advice: the first milestone of an agent RL project is not "push the success rate higher" but "get the cost of a single training sample down to an acceptable level" — if the cost isn't viable, no algorithmic improvement can scale. This maps directly onto the "sample-efficiency awareness" principle on the Anatomy of an RL System page.

INFO

When reading an agent RL paper, check three things first: 1. Is the environment public or gated? 2. Is the "RL" real PPO/GRPO or just RFT? 3. What is the success-rate measurement protocol (how many samples per task, how many seeds)? The answers determine how much of the paper's conclusions you can trust.

Further Reading ​

References ​

  • Yao, S., et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629. (ICLR 2023)
  • Zhao, T., & Eskenazi, M. (2016). Towards End-to-End Learning for Dialog State Tracking and Management using Deep Reinforcement Learning. SIGDIAL 2016. (DQN for dialogue policy)
  • Zeng, A., et al. (2023). AgentTuning: Enabling Generalized Agent Abilities for LLMs. arXiv:2310.12823. (AgentLM)
  • Qin, Y., et al. (2023). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv:2307.16789.
  • Zhou, S., et al. (2023). WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv:2307.13854. (ICLR 2024)
  • Shinn, N., et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. (NeurIPS 2023)
  • Lightman, H., et al. (2023). Let's Verify Step by Step. arXiv:2305.20050. (process reward model)
  • DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948. (large-scale reasoning via RLVR)
  • Brohan, A., et al. (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv:2307.15818.