Transfer a controller, not an animation
In sim-to-real locomotion, engineers design or learn a controller using a simulated body, then test that controller on a physical robot. At each control update it receives observations and produces commands. The physical motors and ground must respond closely enough for the next observation to remain useful.
Consider a left foot lifting while the right foot supports the body. In simulation, the policy has learned how ankle and hip commands affect the next landing. A slippery floor or delayed motor response changes that relationship. Replaying the same joint targets does not recreate the same forces.
For the mechanics of support and foot placement, Read how one walking step is controlled. This guide focuses on transferring the controller that coordinates them.
Build a model that can be checked
| Model choice | Physical question it must answer |
|---|---|
| URDF or MJCF description | Are the links, axes and joint order correct? MJCF also describes engine-specific simulation elements. |
| Joint limits | Which positions and velocities are allowed, and what happens near a stop? |
| Mass and inertia | How much effort changes a link’s translation and rotation? |
| Actuator dynamics | How do targets, gains, torque limits and delay become joint effort? |
| Joint friction | What resistance appears during slow movement and reversal? |
| Ground contact | Where does a foot touch, what load can it support and when does it slip? |
| Sensor model | What are the observation frame, noise, bias, sample interval and delay? |
| Physics timestep | How often are the mechanical state and contact constraints updated? |
Keep visual geometry separate from collision and inertial geometry. A detailed foot mesh can still use an approximate contact shape. For a controller that relies on toe or heel contact, that approximation belongs in the experiment record.
The MuJoCo tutorial makes these fields concrete. Inspect an actual G1 model and its recorded joint response
Train the observation-to-action mapping
An observation might contain body rotation, gyro readings, joint positions and velocities, a requested walking velocity and recent actions. A simulator also knows quantities that a physical robot may not measure. Giving the deployed actor perfect base velocity without an estimator would hide a transfer requirement.
The action space must match the controller architecture. Joint-position offsets followed by PD drives are different from torque commands. An offset vector needs a joint order, a scale and a reference pose. The observation normalization used during training must travel with the policy checkpoint.
Reinforcement learning repeatedly collects trajectories, scores them and updates policy parameters. Velocity tracking can reward useful movement; penalties can discourage extreme joint motion or foot slip. Resets expose different starting states. A fall termination limits an episode, while a timeout truncates a still-valid trajectory. Reward design does not itself prove a safe motion.
For the learning methods themselves, Compare reinforcement learning with imitation. The walking tutorial defines a specific policy and evaluation workflow.
Vary the conditions for a reason
| Variation | Possible mismatch addressed | What to retain in the record |
|---|---|---|
| Link mass or payload | Uncertain body parameters or carried load | Which links changed and whether the range is physically plausible |
| Foot friction | Different shoe and floor interaction | Both the coefficient range and the contact model |
| Observation or command delay | Sampling, networking or motor response lag | Delay units, distribution and where it enters the loop |
| Actuator properties | Different response, damping or available effort | Parameter bounds and the saturation rule |
| Pushes and surface changes | Disturbance recovery and foot placement | Timing, direction and whether a push is a force or a state change |
Randomization asks one controller to work across a distribution of models. It does not measure the physical distribution. Excessively broad ranges can make the task impossible; narrow ones can miss the hardware. Keep held-out conditions that were not used for tuning, and report their outcomes separately.
A walking result with its denominator
Google DeepMind’s 2024 study trained 20-joint ROBOTIS OP3 miniature humanoids in MuJoCo, using dynamics randomization and perturbations. Policies issued joint-position targets at 40 Hz. Physical observations combined proprioception with external motion capture. [3]
| Physical walking measurement | Learned policy | Scripted baseline |
|---|---|---|
| Episodes per controller | 10 | 10 |
| Episode duration | 10 s | 10 s |
| Speed measurement interval | Seconds 2–7 | Seconds 2–7 |
| Mean speed | 0.57 m/s | 0.20 m/s |
| Standard error | 0.003 m/s | 0.005 m/s |
These are measured task means, not product top speeds. The separate turning comparison reran failed attempts until ten successes; the learned policy fell in three of thirteen attempts. The study’s limits include miniature scale, motion capture, joint wear, recalibration and sensitivity to battery charge. It does not establish full-size workplace readiness. [3]
The authors provide Open the authors’ released data and notebooks. Those notebooks were not executed for this guide.
A larger robot raises different model questions
UC Berkeley researchers trained a history-based Transformer controller for Agility Robotics Digit in randomized Isaac Gym environments. Their model approximated closed-chain rods with stiff virtual springs. That is a concrete example of adapting training around an engine’s mechanical representation. [4]
The paper’s comparative ten-run success rates on slopes, steps and unstable ground are simulator evaluations. Physical unstable-plank trials were omitted because of damage risk. A hardware video and a simulated trial table should not be merged into a claimed physical success rate. [4]
For another full-size platform, Read the Atlas hardware analysis. Its specifications are separate from either research result here.
Check the interface before energizing a robot
- Match the policy to the exact robot variant, sensor layout, actuator mode and firmware.
- Check joint order, signs, units, quaternion convention, target offsets and limits with recorded data first.
- Measure the actual observation and command timing. Define handling for stale samples and missed updates.
- Verify available effort and temperature limits through the manufacturer’s documented interface.
- Use supervised, controlled trials with a defined test area, fault response and stop criteria appropriate to the machine.
- Record every attempt, including interventions, aborted runs and recovery, before widening the task.
These are engineering checks, not a complete commissioning or conformity procedure. Simulation can expose faults cheaply, but it cannot verify an untested stop circuit, an incorrect cable connection or a battery fault on a physical unit.
Keep the remaining gap visible
An ideal rigid link does not reproduce structural flex. A friction coefficient does not fully describe a worn sole. Motor heating can reduce available effort after the first few minutes, while sensor bias or calibration drift changes the observations. Network delay and camera processing add timing differences that may not appear in a basic mechanics model.
When the robot diverges, compare aligned logs before changing rewards at random. Determine whether the mismatch is in sensing, the actuator response, contact or the policy’s decision. A good fit on one motion should be tested on another.
Follow the return loop from physical logs to a revised simulation
Sources and verification
- MuJoCo model construction and solver settings ↗Google DeepMind · Read 8 October 2026
- MuJoCo dynamics, actuation and contact ↗Google DeepMind · Read 8 October 2026
- Learning Agile Soccer Skills for a Bipedal Robot with Deep Reinforcement Learning ↗Haarnoja and colleagues / Google DeepMind · Read 8 October 2026
Science Robotics, 2024. Table 1 and supplementary walking protocol distinguish physical measurements from simulation.
- Real-world humanoid locomotion with reinforcement learning ↗Radosavovic and colleagues / UC Berkeley · Read 8 October 2026
Science Robotics, 2024. Comparative ten-run success rates are simulated, not a physical trial denominator.
Article history
Added a sourced engineering guide with version-specific references, practical resources and explicit evidence limits.
Report a correction