Choose the experiment before the training budget
A walking training project needs more than a robot file and a train command. Specify what velocity the robot should follow, which observations it receives, what ends a trial and how the final checkpoint will be compared with a baseline. Keep evaluation conditions separate from the reward used for training.
Use the exact task model
The chosen task is G1JoystickFlatTerrain. It uses Playground’s scene_mjx_feetonly_flat_terrain.xml and its 29-joint G1 physics description, combined with Menagerie assets. This is a different XML model from the introductory CPU example. Playground pins its Menagerie dependency to 1b86ece576591213e2b666ebf59508454200ca97. [1] [2]
Playground code is Apache-2.0 and the referenced Unitree G1 model folder carries BSD-3-Clause terms. Retain the licences if redistributing code or assets. A publicly downloadable model does not imply that every commercial robot variant has the same joints or developer interface. [3] [4]
Before attempting training, Run the smaller G1 loading and logging example. It exercises the interface without needing a learned policy.
Write down the policy’s input contract
| Actor input in the inspected task | Reason to include it |
|---|---|
| Local linear velocity and pelvis gyro | Describe body translation and rotation |
| Projected gravity | Describe orientation relative to gravity |
| Requested planar and yaw velocity | Specify the current motion command |
| 29 joint offsets and 29 joint velocities | Describe limb state relative to the reference pose |
| Previous 29 actions | Expose the previous command to the policy |
| Gait-phase sine and cosine | Provide a periodic phase signal |
The critic receives extra training information, including clean state and contact-related channels. Do not assume the actor receives the same vector. A physical implementation also needs an estimator for body linear velocity; encoders alone do not directly measure the floating body’s translation. [5]
Save the order, units, scaling, noise model and normalization with the checkpoint. A vector of the correct length can still be wrong when two joints are swapped or a rotation convention changes.
Define the action and reward separately
The task has 29 joint-position actions. Each target is the bent-knee reference pose plus 0.5 times its action value. The model follows those targets through its configured position actuators. Physics advances at 0.002 s; the policy updates every 0.02 s, or ten physics steps. These are simulation settings. [5]
| Reward design objective | What to inspect in the recorded rollout |
|---|---|
| Velocity tracking | Difference between requested and measured planar and yaw velocity |
| Upright posture | Whether the torso stays oriented for the task |
| Stable contacts | Foot slip, excessive impact and unwanted collision |
| Smooth commands | Changes in the action vector between control updates |
| Limited joint motion | Joint limits, excessive speed and unnecessary motion |
In this pinned implementation, velocity tracking, orientation, foot behavior and joint-limit terms are active. Action-rate, torque, energy and acceleration weights are zero. To study smoother actions or lower effort, change the relevant terms deliberately and compare a separately identified controller version. Do not describe an inactive term as a protection already supplied by the task. [5]
Collect episodes and update the policy
A training iteration gathers state, action and reward sequences from parallel environments. PPO uses those sequences to update its policy and value model. Resets start new attempts after a termination or time limit. Keep timeout truncation distinct from failures so the learning algorithm and evaluation code receive the right meaning.
The G1 configuration sets 200 million training timesteps and 20 evaluation intervals, with an inherited default of 8,192 parallel training environments. Those are configured budgets, not a measured time to convergence. Changing the batch size to fit memory changes the experiment and should be recorded. [6]
The episode wrapper caps a rollout at 1,000 control steps, or 20 simulated seconds. The environment separately checks an inverted torso, selected inter-leg contacts and non-finite values. A user’s evaluation fall criterion may need to be broader; document it instead of calling every timeout a successful walk. [5]
A source-checked Linux GPU recipe
The upstream setup uses uv and a supported accelerator stack. This recipe selects Python 3.12 and the documented CUDA 12 JAX installation. It is not a CPU-only training recommendation. Check the actual GPU backend before allocating the full training task. Package resolution, drivers and asset downloads still require validation on the machine where it runs. [7]
git clone https://github.com/google-deepmind/mujoco_playground.git
cd mujoco_playground
git checkout d59156e099a8785dd58fbce2e25a6ee1f7b5800c
uv venv --python 3.12
source .venv/bin/activate
uv pip install -U "jax[cuda12]" --index-url https://pypi.org/simple
uv --no-config sync --all-extras
uv --no-config run python -c "import jax; print(jax.default_backend())"
uv --no-config run python -c "from mujoco_playground import locomotion; locomotion.load('G1JoystickFlatTerrain')"The backend check must report a GPU before proceeding with this GPU recipe. The first task load can fetch Menagerie assets. Save the resulting environment lock and package list with the experiment, since pinning a repository alone does not record the installed driver or every resolved package.
uv --no-config run train-jax-ppo --env_name G1JoystickFlatTerrain --impl warp --domain_randomization --seed 42The backend is explicit because this task’s settings and the CLI’s generic default differ. Domain randomization is also explicit; its CLI flag is false by default. Preserve that choice for later comparisons. [8]
Separate randomization from an evaluation test
With the G1 randomizer selected, the inspected code varies floor/foot friction from 0.4 to 1.0, link mass by factors of 0.9 to 1.1, joint friction loss by 0.5 to 2.0, armature by 1.0 to 1.05 and torso mass by an added −1 to +1 kg. Joint-reference offsets also vary. It does not randomize command latency or actuator gains. [9]
Observation noise and pushes are separate environment settings. Its pushes change base velocity; that is not the same as a measured impulse from a physical push device. For evaluation, retain the reset seed, command sequence, surface parameters and disturbance time. Test nominal conditions and held-out variations as separately named groups.
- Freeze a checkpoint before choosing the evaluation outcomes.
- Run the same command sequences and seeds for each compared controller.
- Vary one factor first, such as friction, before combining disturbances.
- Record all falls, early stops and assistance, including unsuccessful attempts.
- Keep the training surfaces and held-out surfaces identified in the results.
Create results only from completed trials
Choose completion and fall criteria before running the evaluation. For a velocity command, a useful tracking metric is the square root of the mean squared velocity error over the stated sample window, in m/s. Decide whether to include startup transients and describe that choice. Aggregate per-trial errors using a declared weighting rule.
| Required measurement | What belongs in it |
|---|---|
| Number of trials | All eligible attempts under the declared protocol |
| Completed trials | Attempts satisfying the full duration and task criteria |
| Falls | Attempts satisfying the predeclared fall rule |
| Average tracking error | A metric, unit and aggregation rule, calculated from logs |
| Tested surface conditions | Named material, friction, terrain and randomization settings |
| Controller version | Checkpoint hash, code revision and resolved configuration |
Download the blank summary table
Download the per-trial CSV template
The summary file separates simulation and physical trials. Its measurement cells are empty because this guide did not perform either walking experiment. Empty fields are not zero failures or zero tracking error. The small joint test in the MuJoCo article is not a walking trial.
The training command does not automatically write this benchmark CSV. Instrument an evaluation loop to record the chosen fields. The upstream playback command produces videos and uses a checkpoint directory; a video alone does not supply a trial denominator. [8]
uv --no-config run train-jax-ppo --env_name G1JoystickFlatTerrain --impl warp --play_only --load_checkpoint_path "$CHECKPOINTS_DIR" --seed 43 --num_videos 4Set CHECKPOINTS_DIR to the actual checkpoints parent directory from your run, containing numeric step directories and its configuration. The inspected loader selects the newest numeric child. Record the selected step and hash; do not compare a moving latest checkpoint without naming which model was evaluated. [8]
A frozen policy still needs a hardware plan
Before physical motion, verify the variant, joint map, estimator, timing, target scale and effort limits against the robot’s documented interface. Define supervision, exclusion space, stop behavior and fault criteria for the actual apparatus. Cutting power can remove a biped’s active balance, so the stop response is a physical design question.
Sources and verification
- G1 feet-only physics model ↗MuJoCo Playground · Read 8 October 2026
- Menagerie dependency pin in MuJoCo Playground ↗MuJoCo Playground · Read 8 October 2026
- Playground code licence ↗MuJoCo Playground · Read 8 October 2026
- Menagerie G1 model and licence at the tested commit ↗Google DeepMind / Unitree · Read 8 October 2026
The model folder is BSD-3-Clause. This is the 29-actuator model, not the base commercial G1 specification.
- G1 joystick task observations, actions and rewards ↗MuJoCo Playground · Read 8 October 2026
- Locomotion PPO training configuration ↗MuJoCo Playground · Read 8 October 2026
- MuJoCo Playground installation at the inspected commit ↗Google DeepMind / MuJoCo Playground · Read 8 October 2026
Commit from 7 October 2026. Installation and training were not executed here.
- JAX PPO training and playback commands ↗MuJoCo Playground · Read 8 October 2026
- G1 model randomization function ↗MuJoCo Playground · Read 8 October 2026
Article history
Added a sourced engineering guide with version-specific references, practical resources and explicit evidence limits.
Report a correction