ROBOTICS FIELD NOTESREVIEW EDITION / 8 October 2026
Guides

How robots learn from actions and rewards

Imitation learning, reinforcement learning and simulation explained through ALOHA and a robotic hand experiment.

Research edition · Sources are linked beside the claims.

What the robot learns

In robot learning, a policy is the rule used to choose an action from available observations. The observation might contain camera pixels, joint angles or a history of sensor readings. An action might request a joint position or a change in gripper pose. The training method determines how that rule is fitted. [1]

Two common routes are imitation learning from recorded behavior and reinforcement learning from reward. Simulation is a place to collect experience and test policies. It can support either route; it is not a third competing learning algorithm.

Learning from recorded behavior

ALOHA records camera images, joint positions and commands during human teleoperation. ACT learns to predict several future joint targets from these observations. Training compares predictions with the recorded commands and adjusts the model to reduce error. The trained model then predicts commands from new observations. [2] [3]

The project includes precise physical tasks such as inserting a battery and opening a small condiment container. These make the learning problem tangible. A slight alignment error can affect the next contact, and an error can accumulate over a sequence. The authors identify this drift as a difficulty for imitation learning. Their reported results concern the tasks and setup in the paper; they are not a universal success rate for household work. [2]

Learning from reward

Reinforcement learning uses interaction and a numerical reward. The objective concerns accumulated reward over time. The reward is part of the task design, and the agent's observations may reveal only part of the underlying state. Training therefore depends on what can be measured, what actions are available and how outcomes are scored. [1]

In an illustrative reaching task, the reward can increase as the gripper gets closer to the target. Training uses the rewards from action sequences to improve the policy. It is not given a demonstrated correct command for every observation. A reward for reaching near an object also leaves grasping and carrying it outside the task unless those outcomes are included. [1]

Moving from simulation to hardware

OpenAI's 2018 in-hand manipulation research trained policies in simulation and tested object reorientation on a physical Shadow Dexterous Hand. During training, the researchers varied physical properties such as friction and varied the object's appearance. That process is commonly called domain randomization. [4]

The purpose was to keep the policy from depending on one exact simulated setup. The reported transfer is evidence for that hand, task and training procedure. It does not establish that every contact interaction can be simulated accurately, or that the same policy will operate an unrelated hand. Hardware validation remains a separate step from a successful simulated episode. [4]

For model assumptions and a physical walking study with its trial count, read the sim-to-real guide.

Read the evaluation before the headline

For an article or technical report, record what data trained the policy, what was held back for testing and which conditions changed. Then examine the success definition. Inserting a part within a time limit, touching the target and completing a whole assembly are different outcomes.

EvidenceQuestion for the reader
Training dataWhich objects and actions were recorded?
Test setWere objects or positions held out?
InterventionsDid a person reset or rescue the attempt?
FailuresWere drops, collisions and incomplete runs counted?

Sources and verification

  1. Spinning Up – Key Concepts in Reinforcement Learning ↗OpenAI · Read 8 October 2026

    Educational definitions of observations, actions, policies and rewards.

  2. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware ↗Tony Zhao, Vikash Kumar, Sergey Levine and Chelsea Finn · Read 8 October 2026

    RSS 2023 project page for ALOHA and ACT. Results apply to the tested setup.

  3. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware ↗Tony Z. Zhao and colleagues · Read 8 October 2026

    2023 v1. Table I and section V-C supply the physical Slot Battery rate and denominator.

  4. Learning Dexterous In-Hand Manipulation ↗OpenAI et al. · Read 8 October 2026

    2018 paper revised in 2019. Specific object reorientation task on a Shadow Dexterous Hand.

Article history

Explained the ACT imitation-learning target and contrasted a demonstration command with a reward signal.

Added a contextual link to the simulation series for the next engineering step.

Report a correction