Appearance
Agents and Dialogue Systems
In a nutshell: this page explains agent RL that treats the LLM as the policy — from dialogue-policy learning in early dialogue systems to tool use in the ReAct paradigm, web navigation, and embodied agents. The core questions: what does the environment interface look like, where does the reward come from, and why does the "task success rate" objective change the shape of RLHF.
1. The Agent Paradigm as an MDP: Environment = API / Browser / World
1.1 Defining the Agent
After 2022, "agent" became the central concept in LLM applications: an LLM wrapped as a looping decision-making unit that repeatedly executes "read state → decide action → take action → observe the outcome":
text
┌──────────────┐ action (call API / click page / issue tool call) ┌───────────────┐
│ LLM agent │ ──────────────────────────────────────────────────► │ Environment │
│ (policy π_θ)│ ◄───────────────────────────────────────────────── │ API/browser/ │
└──────────────┘ observation (tool return / page state / error) │ real world │
└───────────────┘In MDP terms (see the Markov decision process page):
text
State s_t : dialogue history + tool returns + environment observations (page DOM, file system, robot sensors)
Action a_t : generate a passage of text that may contain a tool call (e.g., a JSON-formatted function call)
Reward r_t : task success, tool-call validity, user feedback (sparse!)
Termination : task completed, step budget exhausted, or the agent gives up1.2 Three Key Differences (vs. Traditional RLHF)
| Dimension | RLHF (chat alignment) | Agent RL |
|---|---|---|
| Reward | Human preference (RM score) | Task completion (judged by rules/executor) |
| Environment | The model itself | External APIs / browsers / the physical world |
| Sparsity | One terminal score | Equally sparse, but "success" can be checked automatically |
| State | Autoregressive text | Multimodal, long context, partially observable |
This shift to "reward = task success" is the key to understanding agent RL: when the reward can be verified programmatically, there's no need to train a reward model — and that leads straight to RLVR (see the frontier advances page).
2. Dialogue Policy Learning: RL's Long History in Dialogue Systems
2.1 The "Policy" of Task-Oriented Dialogue Systems
Before LLMs, task-oriented dialogue systems (flight booking, weather lookups) used a pipeline architecture: NLU → dialogue policy → NLG. The dialogue policy decides "what the system should say or do next" — naturally an RL problem:
text
State : dialogue state (user intent, slot-filling status)
Action : system acts (ask for a slot, confirm, query the database, give the final reply)
Reward : +1 for task completion, a small cost per dialogue turn, a heavy penalty for wrong confirmationsZhao & Eskenazi (2016) used DQN to learn dialogue policies, showing that deep RL can learn better "ask–answer–confirm" strategies than hand-written rules. Earlier work (the statistical dialogue systems of the 2010s) had also formalized the problem as a POMDP.
2.2 Historical Lesson: Why Dialogue-Policy RL Never Took Off
- Simulators are hard to build: a user simulator is a fiction, and policies learned against it degrade on real human conversations;
- Sparse, noisy rewards: whether the task succeeded is hard to determine automatically;
- Data scarcity: real dialogue data is limited.
This history hands us a general law: the ceiling of RL = environment fidelity × reward quality. The reason LLM agents revived dialogue RL is precisely that they swapped the "environment" from "user simulators" to "executable APIs and browsers" — rewards can finally be verified automatically.
3. The ReAct Paradigm and RL Fine-Tuning of LLM Agents
3.1 ReAct: Making the LLM Explicitly "Think → Act → Observe"
ReAct (Reasoning + Acting), proposed by Yao et al. (2022), is now the standard skeleton for agents: the LLM alternately outputs "Thought → Action → Observation." This format separates "reasoning" from "acting" explicitly, greatly improving debuggability and success rates.
text
Thought: I need to check the weather in Beijing
Action: weather_api(city="Beijing")
Observation: {"temp": 28, "condition": "sunny"}
Thought: It's sunny, so I can recommend going out
Final: It's sunny in Beijing today, 28°C — great weather for being outdoors.ReAct itself is a prompting technique; it doesn't train the model. What was truly revolutionary is that it provided a unified action format — which paved the way for "training agents with RL."
3.2 RL Fine-Tuning of Agents: AgentLM and Friends
When ReAct prompting falls short (long tasks, complex tools, frequent errors), research turns to RL fine-tuning of the LLM to make its tool calls more reliable. Representative work:
- AgentLM (Zeng et al., 2023): build a large-scale tool-call dataset, do SFT first, then further alignment, improving multi-step tool-call success rates;
- RFT (rejection sampling fine-tuning): sample a large number of answers from the model, filter the "successful" subset with rewards or rules, then SFT on that subset — essentially "offline reward maximization," simpler and more stable than PPO;
- ToolLLM (Qin et al., 2023): expand the tool-call tree with depth-first search, then train tool use with RFT.
3.3 Reward Sources: A Ladder of Executability
Reward design for tool-call RL follows a clear ladder:
| Reward level | What it covers | Characteristics |
|---|---|---|
| Process reward | Is the tool-call syntax valid; was a sensible tool called | Dense; computable automatically |
| Outcome reward | Was the tool return used correctly; is the final answer right | Medium density |
| Task reward | Did the whole multi-step task succeed | Sparse; most truthful |
The key engineering judgment: process rewards (e.g., "this API call was well-formed") let RL quickly learn "how to call," but may produce idle behaviors like "calling the right API to no effect"; task rewards are the most truthful but the sparsest. Industry practice leans toward a mix: process rewards to guide, task rewards to gate.
3.4 The ReAct Family: From Reflection to Multi-Round Improvement
ReAct opened a family of training-free methods that "let the LLM judge itself," complementing RL:
| Method | Mechanism | Relationship to RL |
|---|---|---|
| Reflexion (2023) | After a failure, the agent writes "why it failed" into verbal memory so as to avoid it next round | "Verbal RL": improvement via language instead of gradients |
| Self-Refine (2023) | Generate → self-critique → revise, iterating until satisfied | Training-free "internalization of the reward signal" |
| Tree search (ToT/DFS) | Search multiple paths over the reasoning tree | Isomorphic to AlphaGo's search × learning |
| Agent fine-tuning with RL | Actually update model weights with verifiable rewards | The "gradient version" of the above two |
Practical judgment: for small tasks, Reflexion/Self-Refine are enough (zero cost); turn to RL when "reflection prompts no longer work and the weights must change." This follows the "start simple" principle on the RL design principles page.
3.5 A Complete Agent MDP Example: Tool Calling
Let's write "an LLM calls tools to complete a task" as a trainable MDP and see each component's role:
text
Task: look up a stock price for the user and generate an up/down analysis
State s_t : dialogue history (user instruction + generated thoughts/actions/observations) + tool catalog
Action a_t : generate a passage of text (containing a function call such as stock_price(symbol="AAPL"))
Transition : the tool returns JSON -> appended to the observations
Reward r_t : (terminal) does the final answer contain the correct price ± analysis-quality score
Termination : a Final answer is given / the step budget is exceeded
Training data: real user instructions x tool-call trajectories x terminal success/failureThree agent-specific difficulties this example exposes:
- States are "semi-structured": dialogue history mixes natural language with JSON, beyond the fixed-length state assumption of ordinary RL;
- Actions are "generative": the action space is text space, not an enumerated action set — the "action probabilities" in RL must live on the token distribution;
- Rewards are extremely sparse: only terminal success or failure, with no score for intermediate steps — calling for process rewards or curriculum decomposition.
Together, these three difficulties explain why agent RL is harder than both RLHF and classic RL, and why "get prompting working first → then RFT → finally PPO/GRPO" is the industry's recognized escalation path.
4. RL for Web Navigation and Code Generation
4.1 Web Navigation
Task: given an instruction ("book me a flight from Beijing to Shanghai"), the agent must complete operations on real or simulated web pages (clicking, typing, navigating).
- WebArena / WebArena-Lite (Zhou et al., 2023): self-hosted environments running real web applications, evaluating agents' multi-step web operation skills; they have become the de facto standard in this direction;
- Training recipe: first collect expert trajectories offline for SFT, then RL fine-tune on "whether the final task succeeded";
- Difficulty: web states are partially observable (page content is extremely long; scrolling or clicking is needed to reveal it), so the agent must learn "which operations reveal which information."
4.2 Code Generation
When "writing code" is the task, the reward can be verified directly by an executor:
text
State : problem statement + code generated so far + compilation/test results
Action : the next stretch of code (a token sequence)
Reward : unit-test pass rate / passing hidden tests / code execution resultsTasks with "verifiable rewards" like this make RL unusually clean — no RM training needed; test cases serve as the judge. This is the same thread behind the explosive growth of "RL for reasoning" (math, code) in 2024–2025; see the RLVR section on the frontier advances page for details.
4.3 Training Details for Code-Generation RL
RL fine-tuning for code tasks has four engineering points:
| Point | Practice | Pitfall |
|---|---|---|
| Test-set split | Strict separation of train/validation/test cases | Test-case leakage = score inflation |
| Reward granularity | Unit-test pass rate vs. +1 only if all tests pass | Grade partial passes, but don't let RL learn to "cherry-pick the easy ones" |
| Number of samples | 8–64 answers sampled per problem | Too few samples and nothing is learned; too many and costs explode |
| Compilation environment | Execute code in a sandbox against malicious input | No sandbox = disaster (see the safety note below) |
WARNING
Code-execution safety is a hard boundary: an RL-trained agent executes its own generated code inside the environment. If that code can reach the file system or the network, a malicious prompt mixed into the training data could teach the agent destructive behaviors. In industry, execution must be isolated with Docker sandboxes, timeouts, and resource limits. This is cognate with the "safety guardrails are non-negotiable" principle on the Robotics Control and Sim2Real page.
TIP
Code and math are the "ideal laboratory" for agent RL: rewards are automatically verifiable, environments can be repeated indefinitely, and failures can be precisely attributed. If you're new to agent RL, start with "code completion / unit tests" or "math problem solving" — far easier than open-ended web navigation. For how "verifiable rewards" change RLHF, revisit the RLVR section of LLM Alignment: RLHF in Practice.
5. Embodied Agents and VLA: Agents and Robotics Converge
5.1 VLA: Vision-Language-Action Models
Since 2023, the embodied agent direction has merged the LLM's "general language understanding" with the robot's "action output" into a single model (VLA, Vision-Language-Action):
text
Input : camera images + language instructions
Output : robot actions (joint position deltas / end-effector poses)
Training: massive robot teleoperation data + internet data -> SFT, then RL fine-tuning- RT-2 / RT-2-X (Google DeepMind, 2023): improve VLA generalization with web-scale image-text data;
- OpenVLA (2024): an open-source VLA that the community can fine-tune;
- RL fine-tuning makes VLAs more effective at "learning from failure" — but with real robot data scarce, it still happens mostly in simulation.
5.2 The Connection to Sim2Real
RL for embodied agents and Robotics Control and Sim2Real are two ends of the same thing:
- Robotics side: low-level continuous control (joint torques) is learned with PPO/SAC in simulation;
- Agent side: high-level task planning ("pick up the cup first, then pour the water") is learned with LLM + RL;
- The shared challenge: environment fidelity and reward quality. Whether "learned pouring" in simulation survives a real kitchen depends on the simulator's physics and scene coverage.
6. Environment Engineering: API Wrapping and Reward Sources
6.1 "The Environment Is an API"
Almost all of the engineering weight in agent RL sits on the environment side:
| Component | Engineering task | Common pitfall |
|---|---|---|
| Action space | Wrap environment operations into a unified tool schema (name/params/returns) | Inconsistent schemas cause hallucinated calls |
| Observation returns | Structure tool results (JSON/tables) and cap their length | Overlong returns blow the context |
| Reward judging | A task-success judge (rules/executor/LLM judge) | The judge itself is unreliable |
| Reset and sampling | Environments must be resettable and sampleable in parallel | Stateless environments; Docker isolation needed |
| Error handling | Timeouts, retries, graceful degradation on tool failure | Frequent timeouts during training wreck efficiency |
DANGER
Environment bugs are the number-one invisible killer of agent RL: a tool occasionally returning corrupted JSON, a page element occasionally failing to click — these teach the agent quirks like "retry over and over" or "just give up," while the training curve looks perfectly healthy. This matches "environment bugs silently poison learning" on the common pitfalls page exactly, except agent-environment bugs are even harder to debug (they involve multiple services, networks, and permissions). Pre-launch environment health monitoring and failure-injection testing are mandatory.
6.2 Three Routes to Reward Sources
- Rules/executors (test cases, API response validation) — the most reliable; use automatic verification whenever you can;
- LLM judges (a strong LLM answers "did this response complete the task?") — flexible but gameable; calibration needed;
- Human feedback (the traditional RLHF route) — the most truthful but the most expensive; usually reserved for the final stage.
6.3 Engineering Details of Action-Space Design
In agent RL, the "action space" is a collection of tool calls, and its design quality decides whether training succeeds:
| Design point | Good practice | Bad practice |
|---|---|---|
| Tool granularity | One tool per API, with an explicit JSON Schema for the parameters | Cramming all capabilities into one giant tool |
| Number of tools | Expose 10–50 at a time, dynamically loaded per task | Handing over hundreds at once; the agent freezes with indecision |
| Return format | Structured (JSON/tables) + truncation + error codes | Long unstructured natural-language returns |
| Failure semantics | Tool failures return readable errors instead of throwing | Silent failures (the agent thinks it succeeded) |
| Context management | Summarize historical observations; keep key state | Context balloons without bound; the model loses focus |
The easiest pitfall to fall into is "dirty success/failure signals in tool returns" — the agent learns the bad habit of "calling the tool without looking at the result." If you want a trainable environment, first make tool returns stable, structured, and clearly success/failure-distinguishing. This aligns with the "the environment is the product" thesis on the Anatomy of an RL System page.
7. The Evaluation Problem: Beyond Success Rate
7.1 Why "Task Success Rate" Alone Isn't Enough
Evaluating agents only by "did it succeed in the end" hides a host of problems:
| Dimension | Why it matters | Common distortion |
|---|---|---|
| Step efficiency | Succeeding in 20 steps vs. 5 steps is a different quality | Long detours still count as success |
| Failure modes | Is it "can't call tools" or "misread the task"? | Success rate alone can't attribute blame |
| Generalization | Does it still work on a different task distribution or site layout? | In-distribution score inflation |
| Safety | Did it take dangerous actions (deleting files, placing orders)? | The highest-success agent may be the most dangerous |
| Robustness | How many retries until success? | A one-off success may be luck |
Reporting protocol: task success rate + average steps + failure-mode distribution + multi-seed variance — none can be omitted. This is the protocol from the evaluation and benchmarks page, deployed in the agent setting.
7.2 The Generalization Dilemma
The most painful bottleneck in agent RL right now is generalization: an agent trained on WebArena falls apart on a different website. The reason is that the distribution of web pages and APIs is far too broad for any training distribution to cover. Industry routes forward:
- Larger-scale sampling from diverse environments;
- Few-shot adaptation (show the agent a few examples of the new API before letting it work);
- Write "generalization" into the training objective itself (the domain-randomization idea — see Robotics Control and Sim2Real).
7.3 A Map of Agent Evaluation Benchmarks
You can't evaluate agents only on your own environment; know the mainstream public benchmarks:
| Benchmark | What it tests | Characteristics |
|---|---|---|
| WebArena (2023) | Multi-step web tasks (booking, shopping, admin) | Self-hosted real applications; reproducible tasks |
| SWE-bench (2023) | Fixing code from real GitHub issues | Directly tests "engineering ability"; hard and real |
| GAIA (2023) | General-assistant multi-step tasks | Tests reasoning + tools + multimodality combined |
| ToolBench / API-Bank | Tool-calling ability | Tests API selection and call correctness |
Usage: a paper's "success rate" for an agent must state which benchmark and which protocol were used; when comparing across benchmarks, first check whether the task difficulty is comparable. This also echoes the "benchmark map + reporting standards" methodology on the evaluation and benchmarks page.
8. Frontiers and Criticisms
8.1 Frontier Directions
- Scaling up RLVR: DeepSeek-R1 proved that pure RL (no SFT cold start) can elicit reasoning; agent versions are following;
- Self-improvement loops: agents using feedback to improve their own prompts or policies (Reflexion et al.), naturally complementary to RL's "online sampling";
- Multi-agent agents: multiple LLM agents collaborating on complex tasks — which circles back to the non-stationarity problems of multi-agent RL.
8.2 Criticisms and Cold Water
- Opaque evaluation: many agent papers report success rates on self-built environments, making cross-paper comparison poor;
- "RL training" vs. "prompt engineering": some papers' "RL fine-tuning" is actually just RFT (rejection-sampling SFT), not true policy gradients;
- Cost: every sample in agent RL requires multi-step environment interaction, and the token cost scales with the failure rate — training agents with RL costs an order of magnitude more than alignment.
8.3 Cost Accounting: What Agent RL Actually Costs
Do the math first, then decide whether to adopt agent RL:
text
Cost of one training sample = tokens per task interaction x number of attempts x unit price
Example (web-navigation task, 7B model, internal pricing $0.3 per million tokens):
5,000 tokens per task on average (input + output) x 32 samples = 160k tokens per task
A 10k-task training set -> 1.6B tokens ≈ $480 (sampling only)
Add the training itself (a few epochs each of SFT/PPO) plus failure retries -> multiply by another 2–4x| Stage | Cost source | Control lever |
|---|---|---|
| Data sampling | Multiple rollouts per task | Cap the sample count; validate on a small task set first |
| Environment interaction | Web/API calls, simulators | Cache reusable trajectories; parallelize |
| Training | PPO/GRPO multi-model forward passes | Get it working with a smaller model before scaling up |
| Failure retries | Rework caused by environment instability | Environment health monitoring; failure injection |
Engineering advice: the first milestone of an agent RL project is not "push the success rate higher" but "get the cost of a single training sample down to an acceptable level" — if the cost isn't viable, no algorithmic improvement can scale. This maps directly onto the "sample-efficiency awareness" principle on the Anatomy of an RL System page.
INFO
When reading an agent RL paper, check three things first: 1. Is the environment public or gated? 2. Is the "RL" real PPO/GRPO or just RFT? 3. What is the success-rate measurement protocol (how many samples per task, how many seeds)? The answers determine how much of the paper's conclusions you can trust.
Further Reading
- RLHF and Human-Feedback Alignment — the engine under agent RL: preferences, RM, KL constraints.
- LLM Alignment: RLHF in Practice — the full spectrum from chat alignment to RLVR / reasoning RL.
- Policy gradient methods — implementation details of PPO/GRPO in agent fine-tuning.
- Reward engineering — process vs. outcome rewards; reward design for task success.
- Frontier advances — where ReAct, AgentLM, RFT, WebArena, and DeepSeek-R1 sit.
- Robotics Control and Sim2Real — low-level control and sim-to-real transfer behind embodied-agent VLAs.
References
- Yao, S., et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. arXiv:2210.03629. (ICLR 2023)
- Zhao, T., & Eskenazi, M. (2016). Towards End-to-End Learning for Dialog State Tracking and Management using Deep Reinforcement Learning. SIGDIAL 2016. (DQN for dialogue policy)
- Zeng, A., et al. (2023). AgentTuning: Enabling Generalized Agent Abilities for LLMs. arXiv:2310.12823. (AgentLM)
- Qin, Y., et al. (2023). ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs. arXiv:2307.16789.
- Zhou, S., et al. (2023). WebArena: A Realistic Web Environment for Building Autonomous Agents. arXiv:2307.13854. (ICLR 2024)
- Shinn, N., et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. arXiv:2303.11366. (NeurIPS 2023)
- Lightman, H., et al. (2023). Let's Verify Step by Step. arXiv:2305.20050. (process reward model)
- DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948. (large-scale reasoning via RLVR)
- Brohan, A., et al. (2023). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv:2307.15818.