Appearance
Robot Control and Sim2Real
In one sentence: this page explains why robot RL must be learned in simulation, and how to transfer what was learned there onto real machines — from the simulator ecosystem and the sim-to-real gap to domain randomization, through three real cases (OpenAI's dexterous hand, the ANYmal quadruped, and RMA grasping), and ending with an honest answer to "why most industrial robots don't use RL."
1. The RL Problem Setup for Robotics
1.1 A Battlefield Completely Different from Atari
Robot RL differs fundamentally from game RL. Consider the comparison first:
| Dimension | Atari games | Robot control |
|---|---|---|
| Action space | Discrete (≤18 buttons) | Continuous (joint torques/positions, tens to hundreds of dimensions) |
| State | Pixel frames | Joint angles, angular velocities, IMU, force/tactile sensors |
| Cost of trial and error | Millisecond-scale simulation, nearly free | Real hardware: seconds per step, with possible damage |
| Reward | Game score, naturally available | Must be hand-defined (e.g., "move forward"), often sparse |
| Safety | None | Top priority (human safety, equipment safety) |
| Evaluation | Scores are reproducible | Real-world physics experiments are expensive and hard to reproduce exactly |
Two core tensions: the continuous action space determines the algorithm family (SAC/PPO dominate; see the Actor-Critic family); expensive samples dictate that learning must happen in simulation or from offline data.
1.2 Why "Running RL on a Real Robot" Is Nearly Impossible
Do the arithmetic:
- Deep RL typically needs tens of millions to hundreds of millions of environment steps to learn "a walking quadruped gait" — on-policy algorithms like PPO especially so.
- Every step on a real robot (one control cycle in gait control, roughly 20–50 ms) wears down motors and consumes wall-clock time.
- During trial and error, the policy will inevitably stumble first — wall collisions, falls, and flips can damage the hardware directly. That is part of training, but the real world can't afford this tuition.
So since 2016, the mainstream path for robot RL has been: learn through massive trial and error in a simulator, then transfer the policy to the real robot — Sim-to-Real. Simulators cut the "tuition" from thousands of yuan per attempt on real hardware to nearly zero on a GPU.
TIP
Here is a judgment worth remembering: the hard part of robot RL was never the algorithm — it's "where does the data come from." Simulation provides data, but the data "grows crooked" (sim gap); real data is clean but nearly impossible to obtain in sufficient quantity. The engineering history of robot RL is the history of reconciling this tension. For the systematic treatment of "the environment is the product," see Anatomy of an RL System.
2. Simulators and the Reality Gap
2.1 The Mainstream Simulator Ecosystem
| Simulator | Type | Characteristics | Best suited for |
|---|---|---|---|
| MuJoCo | Continuous-body contact simulation | Fast, numerically stable, the academic default | Research, grasping/manipulation |
| PyBullet | Open-source rigid-body simulation | Free, Python API | Prototyping |
| Isaac Gym / Isaac Lab (NVIDIA) | GPU-parallel rigid-body simulation | Tens of thousands of parallel environments on a single GPU | Large-scale parallel training |
| Brax (Google/JAX) | Pure-JAX differentiable simulation | GPU/TPU parallelism, microsecond-level stepping | Large-scale RL research |
| DM Control (DeepMind) | Classic control benchmarks | Same lineage as MuJoCo | Evaluation benchmarks |
| NVIDIA Omniverse / Isaac Sim | Physics + rendering in one | High-fidelity rendering + physics | Vision-to-manipulation transfer |
Selection advice and installation notes can be found in Datasets & Tools.
2.2 Where the Sim-to-Real Gap Comes From
Differences between simulation and reality (the domain gap) have five main sources:
| Source | Manifestation | Impact |
|---|---|---|
| Dynamics model error | Inaccurate friction coefficients, mass distributions, actuator delays | Mismatched force/torque outputs on the real machine |
| Visual rendering differences | Different materials, lighting, textures, sensor noise | Vision-dependent policies "don't recognize" reality |
| Actuator characteristics | Real motors have bandwidth limits, dead zones, friction, thermal drift | High-frequency actions distort; overheating |
| Sensor noise | Real IMUs/encoders have noise and drift | State-estimation errors get amplified by the policy |
| Non-stationarity | Battery voltage drops, joint wear, ground slippage | The same policy performs inconsistently |
This is the root of "wins in simulation, loses on hardware": the "optimal solution" a policy learns in simulation tends to overfit to simulation-specific details — the robotics version of "environment bugs silently shape learning" from Common Pitfalls & Antipatterns.
2.3 A Taxonomy of Gap-Closing Methods
Industry solutions to the sim-to-real gap fall into four families; understanding the taxonomy matters more than memorizing individual tricks:
| Family | Idea | Representative work | When it fits |
|---|---|---|---|
| Domain randomization | Randomize parameters during training to force out a robust policy | Tobin 2017, OpenAI dexterous hand | Parameter ranges estimable; zero-shot transfer desired |
| System ID / proxy physics | First tune simulator parameters to match reality | Hwangbo 2019 (ANYmal) | Dynamics error dominates; real data collectible |
| Adaptation | Identify the environment online at test time and adjust | RMA 2021, Kumar 2021 | Environment keeps changing; continued stability needed |
| Real-world fine-tuning | Simulation pretraining + real-data fine-tuning | OpenAI Rubik's cube (1% real data) | Real data scarce but obtainable; safety manageable |
Methodological point: the four are not mutually exclusive. Industrial projects often start with system ID to shrink the gap, add domain randomization for robustness, and finish with real-world fine-tuning. Choose by asking "what is the main source of the gap": imprecise parameters → identification/randomization; dynamic environment changes → adaptation; unrealistic rendering → visual randomization.
3. Domain Randomization: the Policy Must Cope With a World It Has Never Seen
3.1 The Core Idea
Domain randomization (Tobin et al., 2017) is disarmingly simple: during training, treat the simulator's physical/visual parameters as random variables, and sample a fresh set at the start of every episode:
text
Sampled randomly at each training reset:
μ_friction ~ U[0.3, 1.0] friction coefficient
m_joint ~ U[0.8, 1.2] × nominal mass
lighting / texture / camera pose randomized visual parameters
......
The learning goal: a policy robust across all of these parameter settingsThe intuition: turn "the simulation is inaccurate" into "the simulation could be anything," forcing the policy to learn robust behavior that doesn't depend on any single parameter value. The trained policy is then deployed directly on the real robot — zero-shot transfer.
3.2 Comparing the Kinds of Domain Randomization
| Type | What is randomized | Representative work | Pros | Limits |
|---|---|---|---|---|
| Physics-parameter randomization | Friction, mass, damping, etc. | Peng 2018, Hwangbo 2019 | No real data needed | Ranges must be set by hand |
| Vision randomization | Textures, lighting, camera pose | Tobin 2017, Bousmalis 2018 | Fixes "doesn't recognize reality" | Limited coverage of real-world variety |
| Automatic domain randomization (ADR) | Auto-expand ranges via "real evaluation + sim generation" | OpenAI 2019 (dexterous hand) | Ranges expand automatically | Needs the real environment for validation |
WARNING
The hidden cost of domain randomization: it moves the burden from "make the simulation accurate" to "choose the randomization ranges." Ranges too small → transfer fails; ranges too large → the policy learns "overly cautious" behavior (imagine a quadruped making every movement timid). Picking the ranges itself requires real-machine iteration — it is not free.
4. Case 1: OpenAI's Dexterous Hand (Dexterous In-Hand Manipulation)
4.1 2018: Rotating the Block
OpenAI (Andrychowicz et al., 2018) trained a Shadow Dexterous Hand — a 24-DoF humanoid hand — to rotate a tri-colored block in its palm to a target orientation. This task counts as dexterous manipulation even for humans, and no robot had learned it before. The recipe:
- 1952 parallel simulated environments (thousands of CPU cores), trained with asynchronous PPO.
- Domain randomization: friction, mass, block size, and textures all randomized.
- This amounted to roughly 100 years of equivalent real-world manipulation experience (equivalent years, in the paper's own words), trained in about 50 hours of simulated time.
- Deployment: the policy transferred zero-shot to the real dexterous hand with about 50% success — versus nearly 100% in simulation.
4.2 2019: The Rubik's Cube
In 2019 OpenAI pushed the same platform toward "solving a Rubik's Cube with one hand." The difficulty ratcheted up a level: the task requires long-horizon memory (which face holds which color). The solution:
- ADR (Automatic Domain Randomization): no more hand-set randomization ranges. Instead, evaluate the policy on the real robot; if it fails, automatically widen the randomization ranges in simulation and retrain — looping until the real machine succeeds.
- LSTM policy network: give the policy a hidden state so it integrates visual information across time, solving the problem of "losing track of the target state while it is occluded."
- Result: with 50% randomized initial states, the real robot solved the cube with about 20% success (and fine-tuning on 1% real-world-collected data pushed the success rate up further).
INFO
The dexterous-hand cases illustrate a key point: sim-to-real is not just a "training trick" problem — it is also a "model architecture" problem. The LSTM memory in the cube task exists precisely to bridge the gap between "simulation assumes complete observation" and "reality loses observation." To understand this "partial observability → memory" mapping, revisit the progression from multi-armed bandits to full RL — state representations keep evolving with the task.
5. Case 2: The ANYmal Quadruped — RL Beats Traditional Controllers
5.1 2019: Learning to Run
The ETH Zürich team (Hwangbo et al., 2019, Science Robotics) trained walking and running gaits on the ANYmal quadruped entirely in simulation, then transferred them zero-shot to the real machine. The key conclusions at the time:
- Proxy physics fine-tuned the dynamics model so that leg–ground contact in simulation better matched reality, combined with randomized training.
- The learned gaits beat the previously hand-designed model-predictive-control (MPC) gaits in both speed and robustness, and adapted to non-ideal terrain such as slopes and gravel.
5.2 2020: Perceptive Locomotion
Lee et al. (2020, Science Robotics) went further on the same platform with perceptive locomotion: using the onboard camera to perceive terrain (rocks, steps, ramps) and cross terrain that had never been seen before. Engineering highlights:
- Terrain was randomly sampled in simulation, including "invisible regions" that were fully simulated during training.
- The trained high-level policy selects gait modes while the low-level controller executes — a hierarchical decomposition that also shows up in autonomous driving's hierarchical decision-making.
- Real-machine demos: stable walking on grass, rubble, snow, and slopes.
5.3 The Algorithm Side: Why PPO
Nearly all of these quadruped works use PPO (or PPO variants), for reasons consistent with the analysis on the policy gradient methods page:
| Requirement | PPO's answer |
|---|---|
| Continuous actions, stochastic policy | Gaussian policy + reparameterized sampling |
| Training stability (thousands of hours of simulation without collapse) | Clipping limits the size of each update |
| Only needs data from the current policy | On-policy; simulation data is cheap, so sample efficiency hardly matters |
| Simple and reliable | Few hyperparameters; the defaults just work |
Switch to SAC only when sample efficiency becomes the priority (e.g., fine-tuning on real data); see why continuous control is SAC's home turf.
6. Case 3: Grasping and Manipulation — RMA and Real-Machine Fine-Tuning
6.1 RMA: Learning the "Adaptation" Itself
In 2021, Kumar et al. proposed RMA (Rapid Motor Adaptation) for a specific problem in the "simulation training → real machine" pipeline: dynamics mismatch varies with the environment.
- Stage one: train a policy network and an "environment encoder" jointly in simulation — the encoder infers "the current environment's dynamics parameter vector" from a short history of observations (about 0.5 seconds of joint/IMU data).
- Stage two: concatenate the encoder's output with the policy's input; at test time the encoder performs online inference from real sensor readings, and the policy adjusts its force output in real time.
Result: the RMA-trained quadruped policy, transferred to a real ANYmal, withstood unknown payloads and friction significantly better than static domain-randomization policies.
6.2 Real-Machine Fine-Tuning: Reinforcing Sim2Real
Another route is "simulation pretraining + real-machine fine-tuning" (Sim2Real + fine-tuning). Typical practice:
- First learn a rough policy in simulation that completes most of the task (avoiding the danger of "starting from scratch" on hardware).
- Then fine-tune for a few episodes on the real robot using a sample-efficient off-policy algorithm (SAC).
- Safety mechanisms are required: action limits, force limits, emergency stops — the positive example of the "safety guardrails" section in Common Pitfalls & Antipatterns.
TIP
The most common engineering mistake is treating "converged in simulation" as "problem solved." The only criterion for a qualified sim-to-real pipeline is real-deployment metrics; simulation numbers are for internal ablation comparisons only. Any "ship it based on simulation curves alone" process should be sent back — consistent with the "don't just watch one metric" spirit of the evaluation & benchmarks page.
6.3 Case 4: The Industrial Reality of Grasping and Manipulation
Grasping (picking) is the manipulation task closest to real deployment, but industry attitudes toward "RL grasping" are deeply split:
| Route | Approach | Status |
|---|---|---|
| Classical visual servoing | Traditional perception + grasp planning | The absolute industrial mainstream (palletizing, sorting, etc.) |
| Deep-RL grasping | Learn policies in simulation, then transfer | Active research, rare deployment (pilots in container and random-item sorting) |
| Deep supervised grasping | Learn "grasp points" from heavily labeled / self-supervised data (not RL) | The actually deployed route at Google, Amazon, and others |
Why RL grasping is hard to deploy: the hard part of grasping is "perceptual geometry" (identifying graspable locations), not "control" — and geometric recognition suits supervised learning far better. RL's value lies in "policy decisions" (which object to grasp, how to sequence the grasps), which only becomes apparent when objects are dynamic, occluded, or require ordering (e.g., unloading, bin clearing). This confirms once more: before RL enters a real system, first ask whether the hard part is perception or decision-making.
6.4 Why Real Data Is So Expensive
Dig one layer deeper: real robot data is "expensive" not only to collect, but also to process:
| Cost item | Specifics | Magnitude vs. simulation |
|---|---|---|
| Collection | Real-machine operation is slow, with human intervention | 10^3–10^4 times the wall-clock |
| Labeling | Human labeling of targets, trajectories, and safety events | Significant |
| Cleaning | Sensor noise; identifying failed episodes | Noticeable |
| Reproducibility | Real-machine results cannot be reproduced exactly | Incomparable (simulation resets) |
| Risk | Hardware damage, human injury | Unquantifiable |
Engineering corollary: in robot RL, real data is not the "training workhorse" — it is "calibration and validation." The correct usage: learn the bulk of the policy in simulation; use real data only to (a) verify the sim gap, (b) fine-tune the last percentage point, and (c) run final acceptance. Trying to train the entire policy from real data is the most common budget black hole in robot RL projects.
7. Reality Check: Why Most Industrial Robots Don't Use RL
7.1 Why Classical Control Still Dominates
Industrial robots (welding, spray-painting, assembly, palletizing) overwhelmingly still run classical control:
- PID / computed-torque control: with known dynamics, it solves fast, is interpretable, and has mathematically guaranteed stability.
- MPC (model predictive control): receding-horizon optimization with constraint satisfaction, handling complex trajectories and constraints.
- Trajectory planning: pre-programmed or teach-pendant demonstrations — deterministic, auditable, easy to debug.
| Dimension | Classical control | Deep-RL control |
|---|---|---|
| Interpretability | High (parameters are physical quantities) | Low (black-box policy) |
| Safety guarantees | Theoretical bounds | Experience and guardrails |
| Data requirements | Model parameter calibration suffices | Massive interaction |
| Generalization to "unseen environments" | Weak | Strong (given enough data) |
| Maintenance | Engineers can repair it | Faults are hard to localize |
7.2 When Is RL Worth Deploying
A usable decision framework (echoing the checklist on the RL Design Principles page):
- The task is hard to hand-program: gaits, grasping, navigating in clutter — the parts where rules can't be written.
- A usable simulator exists: dynamics or vision with enough fidelity; domain randomization can cover the remainder.
- Policy mistakes are affordable — or guardrails can catch them.
- A continuous data feedback loop exists: real data keeps flowing in to improve the policy.
Only when all four hold does RL's cost-effectiveness check pass; otherwise classical control or heuristics are the safer bet.
8. The Algorithm Leads: PPO and SAC
Finally, a wrap-up of the algorithm side. The algorithm spectrum of robot RL is highly concentrated:
text
Continuous-control tasks
├─ on-policy (costly samples but stable): PPO ← the robot mainstream
│ └─ variants: PPO + domain randomization / ADR / RMA encoder
├─ off-policy (sample-efficient): SAC (real-machine fine-tuning) ← TD3's close cousin
└─ model-based (most sample-efficient): Dreamer / TD-MPC2 ← frontier direction| Algorithm | on/off-policy | Sample efficiency | Stability | Robot use cases |
|---|---|---|---|---|
| PPO | on | Low | High | Large-scale simulation training (the mainstream) |
| SAC | off | High | Medium | Real-machine fine-tuning, sample-constrained settings |
| TD3 | off | High | Medium | Deterministic-policy tasks |
| DreamerV3 / TD-MPC2 | model-based | Very high | Medium-high | Data-scarce settings, generalization research (frontier) |
Detailed algorithm comparisons and selection advice: see the Actor-Critic family and choosing frameworks & tools; for the model-based route, see model-based RL.
WARNING
Don't port Atari hyperparameters straight to robots. Atari-tuned PPO learning rates and network widths often either fail to converge or crawl in continuous control. Start from hyperparameters published by robot-learning papers (the ANYmal work, for example, released its full training configuration), then tune step by step following Hyperparameter Tuning & Optimization.
9. Real-Machine Deployment Protocol and Safety Checklist
Whatever the algorithm, real-machine deployment is the last gate of sim-to-real. Engineering teams must follow these protocols:
9.1 The Four Stages of Real-Machine Experiments
text
Stage 1 Simulation verification: the policy converges stably under many randomized configurations
Stage 2 Hardware-in-the-loop (HIL): the real controller runs simulated signals to verify the interface
Stage 3 Constrained real-machine runs: slow speed, small workspace, a safety operator present,
per-trial quotas
Stage 4 Gradual release: loosen limits only after monitored metrics (success rate, force/torque
limit violations) meet their targets9.2 Safety Guardrail Checklist
| Guardrail | Practice | Purpose |
|---|---|---|
| Physical hard limits | Hard limits on joint angle / torque | Prevent the policy from outputting dangerous actions |
| Emergency stop | Manual e-stop button + automatic safety thresholds | Cut power before irreversible errors |
| Action filtering | Low-pass filtering, rate limits | Smooth out high-frequency jitter |
| Monitoring | Real-time logging of force/position/state; automatic degradation on limit breach | Failures can be attributed |
| Fallback policy | Safe-state controller (e.g., zero torque, preset posture) | Automatic takeover when the policy fails |
DANGER
The hidden cost of real-machine experiments: one failed run may force a full system recalibration. A collision can damage sensors and shift dynamics parameters, invalidating all prior simulation calibration. Hence the industry iron rule: "95% confidence in simulation before spending 5% of the budget on the real machine." For a systematic understanding of this cost structure, see "the environment is the product" in Anatomy of an RL System.
9.3 From Paper to Reproduction: A Robot-RL Checklist
When reading a robot RL paper and attempting to reproduce it, use this checklist as a "feasibility health check":
| Check | The paper should provide | If it doesn't |
|---|---|---|
| Simulator & version | MuJoCo/Isaac + version number, model files | Use an open-source alternative environment; treat results as reference only |
| Full reward function | Formula and coefficients for every reward term | Near-impossible to match results when missing |
| Hyperparameter table | Learning rate, γ, GAE λ, clip, network architecture | Dig through the official implementation or the issue tracker |
| Randomization ranges | Randomized parameters and their distributions | Different ranges = a different task |
| Training budget | Environment steps, parallel-env count, wall-clock time | Results not comparable without matched compute |
| Transfer protocol | Real-machine test count, success-rate definition | No real-deployment data = a simulation-only experiment |
Three reality reminders:
- A robot paper's "success rate" depends heavily on environment details (initial randomization, reward shaping); a 20–30 percentage point gap between implementations is normal.
- Prefer open-source implementations over writing from scratch — get the official code running before changing anything. This is the positive posture of the "failed paper reproduction diagnosis checklist" in Common Pitfalls & Antipatterns.
- Only real-machine data can arbitrate the gap between simulation results and real deployment — for a paper's "sim-to-real success," also check what hardware and what task difficulty it actually transferred to.
Further Reading
- The Actor-Critic Family — the SAC/PPO/TD3 lineage and continuous-control selection.
- Policy Gradient Methods — the mechanics and implementation of PPO's clipped objective.
- Model-Based RL and World Models — the "learn dynamics first, then plan" route on robots, plus the model-error trap.
- Choosing Frameworks & Tools — trade-offs among Isaac Gym / Brax / SB3 for robot tasks.
- Datasets & Tools — a list of simulators including MuJoCo, Isaac, and DM Control.
- Common Pitfalls & Antipatterns — concrete forms of safety guardrails and sample-efficiency illusions in robot engineering.
References
- Tobin, J., et al. (2017). Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World. arXiv:1703.06907.
- Peng, X. B., et al. (2018). Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. ICRA 2018. (arXiv:1805.09932)
- Bousmalis, K., et al. (2018). Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping. ICRA 2018. (arXiv:1710.06537)
- Andrychowicz, M., et al. (2020). Learning Dexterous In-Hand Manipulation. IJRR 39(1), 3–20. (arXiv:1808.00177)
- OpenAI, et al. (2019). Solving Rubik's Cube with a Robot Hand. arXiv:1910.07113.
- Hwangbo, J., et al. (2019). Learning Locomotion Using Deep Reinforcement Learning. Science Robotics, 4(29), eaau5872.
- Lee, J., et al. (2020). Learning Quadrupedal Locomotion over Challenging Terrain. Science Robotics, 5(47), eabc5986.
- Kumar, A., et al. (2021). RMA: Rapid Motor Adaptation for Legged Robots. arXiv:2109.06113. (RSS 2021)
- Haarnoja, T., et al. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. arXiv:1801.01290.
- Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347.
- Todorov, E., Erez, T., & Tassa, Y. (2012). MuJoCo: A Physics Engine for Model-Based Control. IROS 2012.