Skip to content

Robot Control and Sim2Real

On this page Real-world physical trial-and-error is too expensive, so learning in simulation has become the mainstream — the simulator ecosystem (MuJoCo, Isaac), domain randomization, the Sim2Real gap and how to close it; case studies from OpenAI's dexterous hand and the ANYmal quadruped to robotic grasping.

Robot Control and Sim2Real ​

In one sentence: this page explains why robot RL must be learned in simulation, and how to transfer what was learned there onto real machines — from the simulator ecosystem and the sim-to-real gap to domain randomization, through three real cases (OpenAI's dexterous hand, the ANYmal quadruped, and RMA grasping), and ending with an honest answer to "why most industrial robots don't use RL."

1. The RL Problem Setup for Robotics ​

1.1 A Battlefield Completely Different from Atari ​

Robot RL differs fundamentally from game RL. Consider the comparison first:

DimensionAtari gamesRobot control
Action spaceDiscrete (≤18 buttons)Continuous (joint torques/positions, tens to hundreds of dimensions)
StatePixel framesJoint angles, angular velocities, IMU, force/tactile sensors
Cost of trial and errorMillisecond-scale simulation, nearly freeReal hardware: seconds per step, with possible damage
RewardGame score, naturally availableMust be hand-defined (e.g., "move forward"), often sparse
SafetyNoneTop priority (human safety, equipment safety)
EvaluationScores are reproducibleReal-world physics experiments are expensive and hard to reproduce exactly

Two core tensions: the continuous action space determines the algorithm family (SAC/PPO dominate; see the Actor-Critic family); expensive samples dictate that learning must happen in simulation or from offline data.

1.2 Why "Running RL on a Real Robot" Is Nearly Impossible ​

Do the arithmetic:

  • Deep RL typically needs tens of millions to hundreds of millions of environment steps to learn "a walking quadruped gait" — on-policy algorithms like PPO especially so.
  • Every step on a real robot (one control cycle in gait control, roughly 20–50 ms) wears down motors and consumes wall-clock time.
  • During trial and error, the policy will inevitably stumble first — wall collisions, falls, and flips can damage the hardware directly. That is part of training, but the real world can't afford this tuition.

So since 2016, the mainstream path for robot RL has been: learn through massive trial and error in a simulator, then transfer the policy to the real robot — Sim-to-Real. Simulators cut the "tuition" from thousands of yuan per attempt on real hardware to nearly zero on a GPU.

TIP

Here is a judgment worth remembering: the hard part of robot RL was never the algorithm — it's "where does the data come from." Simulation provides data, but the data "grows crooked" (sim gap); real data is clean but nearly impossible to obtain in sufficient quantity. The engineering history of robot RL is the history of reconciling this tension. For the systematic treatment of "the environment is the product," see Anatomy of an RL System.

2. Simulators and the Reality Gap ​

2.1 The Mainstream Simulator Ecosystem ​

SimulatorTypeCharacteristicsBest suited for
MuJoCoContinuous-body contact simulationFast, numerically stable, the academic defaultResearch, grasping/manipulation
PyBulletOpen-source rigid-body simulationFree, Python APIPrototyping
Isaac Gym / Isaac Lab (NVIDIA)GPU-parallel rigid-body simulationTens of thousands of parallel environments on a single GPULarge-scale parallel training
Brax (Google/JAX)Pure-JAX differentiable simulationGPU/TPU parallelism, microsecond-level steppingLarge-scale RL research
DM Control (DeepMind)Classic control benchmarksSame lineage as MuJoCoEvaluation benchmarks
NVIDIA Omniverse / Isaac SimPhysics + rendering in oneHigh-fidelity rendering + physicsVision-to-manipulation transfer

Selection advice and installation notes can be found in Datasets & Tools.

2.2 Where the Sim-to-Real Gap Comes From ​

Differences between simulation and reality (the domain gap) have five main sources:

SourceManifestationImpact
Dynamics model errorInaccurate friction coefficients, mass distributions, actuator delaysMismatched force/torque outputs on the real machine
Visual rendering differencesDifferent materials, lighting, textures, sensor noiseVision-dependent policies "don't recognize" reality
Actuator characteristicsReal motors have bandwidth limits, dead zones, friction, thermal driftHigh-frequency actions distort; overheating
Sensor noiseReal IMUs/encoders have noise and driftState-estimation errors get amplified by the policy
Non-stationarityBattery voltage drops, joint wear, ground slippageThe same policy performs inconsistently

This is the root of "wins in simulation, loses on hardware": the "optimal solution" a policy learns in simulation tends to overfit to simulation-specific details — the robotics version of "environment bugs silently shape learning" from Common Pitfalls & Antipatterns.

2.3 A Taxonomy of Gap-Closing Methods ​

Industry solutions to the sim-to-real gap fall into four families; understanding the taxonomy matters more than memorizing individual tricks:

FamilyIdeaRepresentative workWhen it fits
Domain randomizationRandomize parameters during training to force out a robust policyTobin 2017, OpenAI dexterous handParameter ranges estimable; zero-shot transfer desired
System ID / proxy physicsFirst tune simulator parameters to match realityHwangbo 2019 (ANYmal)Dynamics error dominates; real data collectible
AdaptationIdentify the environment online at test time and adjustRMA 2021, Kumar 2021Environment keeps changing; continued stability needed
Real-world fine-tuningSimulation pretraining + real-data fine-tuningOpenAI Rubik's cube (1% real data)Real data scarce but obtainable; safety manageable

Methodological point: the four are not mutually exclusive. Industrial projects often start with system ID to shrink the gap, add domain randomization for robustness, and finish with real-world fine-tuning. Choose by asking "what is the main source of the gap": imprecise parameters → identification/randomization; dynamic environment changes → adaptation; unrealistic rendering → visual randomization.

3. Domain Randomization: the Policy Must Cope With a World It Has Never Seen ​

3.1 The Core Idea ​

Domain randomization (Tobin et al., 2017) is disarmingly simple: during training, treat the simulator's physical/visual parameters as random variables, and sample a fresh set at the start of every episode:

text
Sampled randomly at each training reset:
  μ_friction  ~ U[0.3, 1.0]                       friction coefficient
  m_joint     ~ U[0.8, 1.2] × nominal             mass
  lighting / texture / camera pose randomized     visual parameters
  ......
The learning goal: a policy robust across all of these parameter settings

The intuition: turn "the simulation is inaccurate" into "the simulation could be anything," forcing the policy to learn robust behavior that doesn't depend on any single parameter value. The trained policy is then deployed directly on the real robot — zero-shot transfer.

3.2 Comparing the Kinds of Domain Randomization ​

TypeWhat is randomizedRepresentative workProsLimits
Physics-parameter randomizationFriction, mass, damping, etc.Peng 2018, Hwangbo 2019No real data neededRanges must be set by hand
Vision randomizationTextures, lighting, camera poseTobin 2017, Bousmalis 2018Fixes "doesn't recognize reality"Limited coverage of real-world variety
Automatic domain randomization (ADR)Auto-expand ranges via "real evaluation + sim generation"OpenAI 2019 (dexterous hand)Ranges expand automaticallyNeeds the real environment for validation

WARNING

The hidden cost of domain randomization: it moves the burden from "make the simulation accurate" to "choose the randomization ranges." Ranges too small → transfer fails; ranges too large → the policy learns "overly cautious" behavior (imagine a quadruped making every movement timid). Picking the ranges itself requires real-machine iteration — it is not free.

4. Case 1: OpenAI's Dexterous Hand (Dexterous In-Hand Manipulation) ​

4.1 2018: Rotating the Block ​

OpenAI (Andrychowicz et al., 2018) trained a Shadow Dexterous Hand — a 24-DoF humanoid hand — to rotate a tri-colored block in its palm to a target orientation. This task counts as dexterous manipulation even for humans, and no robot had learned it before. The recipe:

  • 1952 parallel simulated environments (thousands of CPU cores), trained with asynchronous PPO.
  • Domain randomization: friction, mass, block size, and textures all randomized.
  • This amounted to roughly 100 years of equivalent real-world manipulation experience (equivalent years, in the paper's own words), trained in about 50 hours of simulated time.
  • Deployment: the policy transferred zero-shot to the real dexterous hand with about 50% success — versus nearly 100% in simulation.

4.2 2019: The Rubik's Cube ​

In 2019 OpenAI pushed the same platform toward "solving a Rubik's Cube with one hand." The difficulty ratcheted up a level: the task requires long-horizon memory (which face holds which color). The solution:

  • ADR (Automatic Domain Randomization): no more hand-set randomization ranges. Instead, evaluate the policy on the real robot; if it fails, automatically widen the randomization ranges in simulation and retrain — looping until the real machine succeeds.
  • LSTM policy network: give the policy a hidden state so it integrates visual information across time, solving the problem of "losing track of the target state while it is occluded."
  • Result: with 50% randomized initial states, the real robot solved the cube with about 20% success (and fine-tuning on 1% real-world-collected data pushed the success rate up further).

INFO

The dexterous-hand cases illustrate a key point: sim-to-real is not just a "training trick" problem — it is also a "model architecture" problem. The LSTM memory in the cube task exists precisely to bridge the gap between "simulation assumes complete observation" and "reality loses observation." To understand this "partial observability → memory" mapping, revisit the progression from multi-armed bandits to full RL — state representations keep evolving with the task.

5. Case 2: The ANYmal Quadruped — RL Beats Traditional Controllers ​

5.1 2019: Learning to Run ​

The ETH Zürich team (Hwangbo et al., 2019, Science Robotics) trained walking and running gaits on the ANYmal quadruped entirely in simulation, then transferred them zero-shot to the real machine. The key conclusions at the time:

  • Proxy physics fine-tuned the dynamics model so that leg–ground contact in simulation better matched reality, combined with randomized training.
  • The learned gaits beat the previously hand-designed model-predictive-control (MPC) gaits in both speed and robustness, and adapted to non-ideal terrain such as slopes and gravel.

5.2 2020: Perceptive Locomotion ​

Lee et al. (2020, Science Robotics) went further on the same platform with perceptive locomotion: using the onboard camera to perceive terrain (rocks, steps, ramps) and cross terrain that had never been seen before. Engineering highlights:

  • Terrain was randomly sampled in simulation, including "invisible regions" that were fully simulated during training.
  • The trained high-level policy selects gait modes while the low-level controller executes — a hierarchical decomposition that also shows up in autonomous driving's hierarchical decision-making.
  • Real-machine demos: stable walking on grass, rubble, snow, and slopes.

5.3 The Algorithm Side: Why PPO ​

Nearly all of these quadruped works use PPO (or PPO variants), for reasons consistent with the analysis on the policy gradient methods page:

RequirementPPO's answer
Continuous actions, stochastic policyGaussian policy + reparameterized sampling
Training stability (thousands of hours of simulation without collapse)Clipping limits the size of each update
Only needs data from the current policyOn-policy; simulation data is cheap, so sample efficiency hardly matters
Simple and reliableFew hyperparameters; the defaults just work

Switch to SAC only when sample efficiency becomes the priority (e.g., fine-tuning on real data); see why continuous control is SAC's home turf.

6. Case 3: Grasping and Manipulation — RMA and Real-Machine Fine-Tuning ​

6.1 RMA: Learning the "Adaptation" Itself ​

In 2021, Kumar et al. proposed RMA (Rapid Motor Adaptation) for a specific problem in the "simulation training → real machine" pipeline: dynamics mismatch varies with the environment.

  • Stage one: train a policy network and an "environment encoder" jointly in simulation — the encoder infers "the current environment's dynamics parameter vector" from a short history of observations (about 0.5 seconds of joint/IMU data).
  • Stage two: concatenate the encoder's output with the policy's input; at test time the encoder performs online inference from real sensor readings, and the policy adjusts its force output in real time.

Result: the RMA-trained quadruped policy, transferred to a real ANYmal, withstood unknown payloads and friction significantly better than static domain-randomization policies.

6.2 Real-Machine Fine-Tuning: Reinforcing Sim2Real ​

Another route is "simulation pretraining + real-machine fine-tuning" (Sim2Real + fine-tuning). Typical practice:

  • First learn a rough policy in simulation that completes most of the task (avoiding the danger of "starting from scratch" on hardware).
  • Then fine-tune for a few episodes on the real robot using a sample-efficient off-policy algorithm (SAC).
  • Safety mechanisms are required: action limits, force limits, emergency stops — the positive example of the "safety guardrails" section in Common Pitfalls & Antipatterns.

TIP

The most common engineering mistake is treating "converged in simulation" as "problem solved." The only criterion for a qualified sim-to-real pipeline is real-deployment metrics; simulation numbers are for internal ablation comparisons only. Any "ship it based on simulation curves alone" process should be sent back — consistent with the "don't just watch one metric" spirit of the evaluation & benchmarks page.

6.3 Case 4: The Industrial Reality of Grasping and Manipulation ​

Grasping (picking) is the manipulation task closest to real deployment, but industry attitudes toward "RL grasping" are deeply split:

RouteApproachStatus
Classical visual servoingTraditional perception + grasp planningThe absolute industrial mainstream (palletizing, sorting, etc.)
Deep-RL graspingLearn policies in simulation, then transferActive research, rare deployment (pilots in container and random-item sorting)
Deep supervised graspingLearn "grasp points" from heavily labeled / self-supervised data (not RL)The actually deployed route at Google, Amazon, and others

Why RL grasping is hard to deploy: the hard part of grasping is "perceptual geometry" (identifying graspable locations), not "control" — and geometric recognition suits supervised learning far better. RL's value lies in "policy decisions" (which object to grasp, how to sequence the grasps), which only becomes apparent when objects are dynamic, occluded, or require ordering (e.g., unloading, bin clearing). This confirms once more: before RL enters a real system, first ask whether the hard part is perception or decision-making.

6.4 Why Real Data Is So Expensive ​

Dig one layer deeper: real robot data is "expensive" not only to collect, but also to process:

Cost itemSpecificsMagnitude vs. simulation
CollectionReal-machine operation is slow, with human intervention10^3–10^4 times the wall-clock
LabelingHuman labeling of targets, trajectories, and safety eventsSignificant
CleaningSensor noise; identifying failed episodesNoticeable
ReproducibilityReal-machine results cannot be reproduced exactlyIncomparable (simulation resets)
RiskHardware damage, human injuryUnquantifiable

Engineering corollary: in robot RL, real data is not the "training workhorse" — it is "calibration and validation." The correct usage: learn the bulk of the policy in simulation; use real data only to (a) verify the sim gap, (b) fine-tune the last percentage point, and (c) run final acceptance. Trying to train the entire policy from real data is the most common budget black hole in robot RL projects.

7. Reality Check: Why Most Industrial Robots Don't Use RL ​

7.1 Why Classical Control Still Dominates ​

Industrial robots (welding, spray-painting, assembly, palletizing) overwhelmingly still run classical control:

  • PID / computed-torque control: with known dynamics, it solves fast, is interpretable, and has mathematically guaranteed stability.
  • MPC (model predictive control): receding-horizon optimization with constraint satisfaction, handling complex trajectories and constraints.
  • Trajectory planning: pre-programmed or teach-pendant demonstrations — deterministic, auditable, easy to debug.
DimensionClassical controlDeep-RL control
InterpretabilityHigh (parameters are physical quantities)Low (black-box policy)
Safety guaranteesTheoretical boundsExperience and guardrails
Data requirementsModel parameter calibration sufficesMassive interaction
Generalization to "unseen environments"WeakStrong (given enough data)
MaintenanceEngineers can repair itFaults are hard to localize

7.2 When Is RL Worth Deploying ​

A usable decision framework (echoing the checklist on the RL Design Principles page):

  1. The task is hard to hand-program: gaits, grasping, navigating in clutter — the parts where rules can't be written.
  2. A usable simulator exists: dynamics or vision with enough fidelity; domain randomization can cover the remainder.
  3. Policy mistakes are affordable — or guardrails can catch them.
  4. A continuous data feedback loop exists: real data keeps flowing in to improve the policy.

Only when all four hold does RL's cost-effectiveness check pass; otherwise classical control or heuristics are the safer bet.

8. The Algorithm Leads: PPO and SAC ​

Finally, a wrap-up of the algorithm side. The algorithm spectrum of robot RL is highly concentrated:

text
Continuous-control tasks
 ├─ on-policy (costly samples but stable): PPO ← the robot mainstream
 │    └─ variants: PPO + domain randomization / ADR / RMA encoder
 ├─ off-policy (sample-efficient): SAC (real-machine fine-tuning) ← TD3's close cousin
 └─ model-based (most sample-efficient): Dreamer / TD-MPC2 ← frontier direction
Algorithmon/off-policySample efficiencyStabilityRobot use cases
PPOonLowHighLarge-scale simulation training (the mainstream)
SACoffHighMediumReal-machine fine-tuning, sample-constrained settings
TD3offHighMediumDeterministic-policy tasks
DreamerV3 / TD-MPC2model-basedVery highMedium-highData-scarce settings, generalization research (frontier)

Detailed algorithm comparisons and selection advice: see the Actor-Critic family and choosing frameworks & tools; for the model-based route, see model-based RL.

WARNING

Don't port Atari hyperparameters straight to robots. Atari-tuned PPO learning rates and network widths often either fail to converge or crawl in continuous control. Start from hyperparameters published by robot-learning papers (the ANYmal work, for example, released its full training configuration), then tune step by step following Hyperparameter Tuning & Optimization.

9. Real-Machine Deployment Protocol and Safety Checklist ​

Whatever the algorithm, real-machine deployment is the last gate of sim-to-real. Engineering teams must follow these protocols:

9.1 The Four Stages of Real-Machine Experiments ​

text
Stage 1  Simulation verification: the policy converges stably under many randomized configurations
Stage 2  Hardware-in-the-loop (HIL): the real controller runs simulated signals to verify the interface
Stage 3  Constrained real-machine runs: slow speed, small workspace, a safety operator present,
         per-trial quotas
Stage 4  Gradual release: loosen limits only after monitored metrics (success rate, force/torque
         limit violations) meet their targets

9.2 Safety Guardrail Checklist ​

GuardrailPracticePurpose
Physical hard limitsHard limits on joint angle / torquePrevent the policy from outputting dangerous actions
Emergency stopManual e-stop button + automatic safety thresholdsCut power before irreversible errors
Action filteringLow-pass filtering, rate limitsSmooth out high-frequency jitter
MonitoringReal-time logging of force/position/state; automatic degradation on limit breachFailures can be attributed
Fallback policySafe-state controller (e.g., zero torque, preset posture)Automatic takeover when the policy fails

DANGER

The hidden cost of real-machine experiments: one failed run may force a full system recalibration. A collision can damage sensors and shift dynamics parameters, invalidating all prior simulation calibration. Hence the industry iron rule: "95% confidence in simulation before spending 5% of the budget on the real machine." For a systematic understanding of this cost structure, see "the environment is the product" in Anatomy of an RL System.

9.3 From Paper to Reproduction: A Robot-RL Checklist ​

When reading a robot RL paper and attempting to reproduce it, use this checklist as a "feasibility health check":

CheckThe paper should provideIf it doesn't
Simulator & versionMuJoCo/Isaac + version number, model filesUse an open-source alternative environment; treat results as reference only
Full reward functionFormula and coefficients for every reward termNear-impossible to match results when missing
Hyperparameter tableLearning rate, γ, GAE λ, clip, network architectureDig through the official implementation or the issue tracker
Randomization rangesRandomized parameters and their distributionsDifferent ranges = a different task
Training budgetEnvironment steps, parallel-env count, wall-clock timeResults not comparable without matched compute
Transfer protocolReal-machine test count, success-rate definitionNo real-deployment data = a simulation-only experiment

Three reality reminders:

  1. A robot paper's "success rate" depends heavily on environment details (initial randomization, reward shaping); a 20–30 percentage point gap between implementations is normal.
  2. Prefer open-source implementations over writing from scratch — get the official code running before changing anything. This is the positive posture of the "failed paper reproduction diagnosis checklist" in Common Pitfalls & Antipatterns.
  3. Only real-machine data can arbitrate the gap between simulation results and real deployment — for a paper's "sim-to-real success," also check what hardware and what task difficulty it actually transferred to.

Further Reading ​

References ​

  • Tobin, J., et al. (2017). Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World. arXiv:1703.06907.
  • Peng, X. B., et al. (2018). Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. ICRA 2018. (arXiv:1805.09932)
  • Bousmalis, K., et al. (2018). Using Simulation and Domain Adaptation to Improve Efficiency of Deep Robotic Grasping. ICRA 2018. (arXiv:1710.06537)
  • Andrychowicz, M., et al. (2020). Learning Dexterous In-Hand Manipulation. IJRR 39(1), 3–20. (arXiv:1808.00177)
  • OpenAI, et al. (2019). Solving Rubik's Cube with a Robot Hand. arXiv:1910.07113.
  • Hwangbo, J., et al. (2019). Learning Locomotion Using Deep Reinforcement Learning. Science Robotics, 4(29), eaau5872.
  • Lee, J., et al. (2020). Learning Quadrupedal Locomotion over Challenging Terrain. Science Robotics, 5(47), eabc5986.
  • Kumar, A., et al. (2021). RMA: Rapid Motor Adaptation for Legged Robots. arXiv:2109.06113. (RSS 2021)
  • Haarnoja, T., et al. (2018). Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. arXiv:1801.01290.
  • Schulman, J., et al. (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347.
  • Todorov, E., Erez, T., & Tassa, Y. (2012). MuJoCo: A Physics Engine for Model-Based Control. IROS 2012.