From sensors to movement
Humanoid robots use cameras and range sensors to estimate the objects, surfaces and free space around them. Software associates those measurements with positions and timestamps, then passes a representation to the robot controller. A bottle label in an image, the bottle’s position in metres and the instruction to pick it up are three different pieces of information.
A useful outline is sensor capture, data processing, object detection, depth estimation, scene representation and robot control. Several stages run together. Depth can arrive directly from a sensor while a detector processes a colour image. New measurements return during movement, so the scene estimate changes as the robot turns its head or reaches.
What a camera measures
An RGB camera records red, green and blue image channels. Stereo vision matches the same scene point in two images. With the cameras’ positions, orientations and projection parameters known, each matched pixel supplies a viewing ray. Their intersection estimates the point’s location in space. This calculation is called triangulation. Real measurements require fitting an estimate because noise can keep the two rays from meeting. [1]
Active depth sensing supplies its own illumination. A time-of-flight device derives range from the travel time or phase shift of returning light. LiDAR is laser ranging, and some LiDAR sensors also produce a dense depth image. The RealSense L515 is one example. Sensor categories therefore overlap. A depth image is a grid of distances; a point cloud places those observations in three-dimensional coordinates. [2]
Naming an object is one calculation
Object detection assigns a category and an image region, often a rectangular box. Semantic segmentation labels pixels by category, such as floor or bottle. Instance segmentation separates individual objects within one category. Mask R-CNN provides an identifiable example of a network that predicts both boxes and individual object masks. None of those outputs, by itself, supplies the full position and orientation needed for a grasp. [3]
Pose estimation adds geometry. A six-dimensional rigid-object pose contains three position coordinates and three rotational degrees of freedom. FoundationPose, published at CVPR 2024, estimates and tracks poses using colour images paired with depth measurements, often called RGB-D. It requires either an object CAD model or reference views, and a separate detector supplies the object region. The paper reports failures involving severe occlusion, little texture and few visible edges. A detected object can still have the wrong orientation. [4]
Hand-eye calibration estimates the rigid relationship between a camera and a robot link. [5]
Coordinate-frame software then keeps track of those relationships while the robot moves. ROS 2’s tf2 library stores a tree of frames with a time history. A head-camera observation can be expressed relative to the robot base or gripper using the relationships valid at capture time. Pairing an old image with the current head pose can place the target in the wrong location. That consequence follows from mixing measurements from different times. [6]
The robot also moves
Visual odometry estimates camera motion from successive images. SLAM, simultaneous localization and mapping, estimates a trajectory while maintaining a map that the system can revisit. Matching a known place can correct accumulated drift. Without this separation between camera motion and scene motion, a stationary bottle can appear to move whenever the robot’s head turns. [7]
ORB-SLAM3 combines visual observations with inertial measurements in its visual-inertial modes. Its map system can start another map after tracking loss and later join maps when it recognizes shared places. The software supports several camera configurations. It is an example of localization machinery, with no assumption that a particular commercial humanoid runs it. [7]
Three-dimensional reconstruction combines observations into surfaces, meshes or occupied volumes. OctoMap stores probabilities in a tree of spatial cells and distinguishes occupied, free and unobserved space. A measured wall supplies evidence for its surface and the clear path of the sensor ray. Space behind that wall remains unobserved. This distinction matters when an arm planner checks a route around furniture. [8]
A documented Atlas architecture
Boston Dynamics describes the hydraulic Atlas parkour system using a time-of-flight camera that produced point clouds at 15 frames per second. Software separated planar surfaces, fitted obstacle faces and tracked their positions. The robot received an approximate course map containing obstacle templates and annotated actions. Live measurements supplied the geometry for its next footholds. [9]
The company gives an example of moving a box 0.5 m sideways. Atlas could find its changed position and adjust the jump. Moving it beyond the search region caused the system to stop. This account describes a prepared parkour course with prior information. Its sensing specifications belong to that historical platform. [9]
Where the estimate breaks
| Condition | What can go wrong |
|---|---|
| Low light | RGB detail becomes weak. Active infrared depth can continue under conditions that defeat a colour image. |
| Reflections | A smooth surface can redirect the emitted beam away from its receiver. |
| Transparent objects | Light can pass through glass and return from the background. |
| Fast movement | Exposure blur or averaging across frames can smear edges and old positions. |
| Bright sunlight | The indoor L515 can lose depth values when ambient infrared light overwhelms its returning signal. |
Occlusion hides the evidence needed to distinguish two objects or estimate a pose. FoundationPose’s documented failure example combines occlusion with weak visual cues. Maintaining a track through a hidden interval means predicting from older observations; it does not create a fresh measurement. [4]
The Stanford stereo notes describe reconstruction errors from noisy observations and imperfect camera calibration. Errors in camera-to-arm geometry introduce another possible source of misplaced reaches. [1]
A map still needs a task
Spatial reasoning can use relations such as bottle above table, handle facing outward or target behind an obstacle. These relations support choices about viewpoint and approach. They do not establish the contents of a closed container or the stability of an unseen surface. A system needs to retain uncertainty instead of treating every inferred property as a measurement.
The next step is described in how humanoid robots turn an instruction into an action. That article follows a bottle from identification through grasping and result verification.
Sources and verification
- CS231A Stereo Systems and Structure from Motion ↗Kenji Hata and Silvio Savarese, Stanford University · Read 8 October 2026
Course notes on triangulation from calibrated viewpoints, observation noise and reconstruction errors.
- RealSense LiDAR Camera L515 user guide ↗Intel RealSense · Read 8 October 2026
July 2020 revision 1.0. Pages 18–20 cover reflections, transparent objects and ambient infrared interference. Sensor physics is used here, with no claim of installation on Atlas.
- Mask R-CNN ↗Kaiming He, Georgia Gkioxari, Piotr Dollár and Ross Girshick · Read 8 October 2026
Original paper distinguishes detection, semantic segmentation and instance segmentation. No benchmark score is generalized to a robot.
- FoundationPose Unified 6D Pose Estimation and Tracking of Novel Objects ↗Bowen Wen, Wei Yang, Jan Kautz and Stan Birchfield · Read 8 October 2026
CVPR 2024 paper. Sections 3 and 5.4 explain RGB-D pose estimation, external detection and failure under combined occlusion and weak visual cues.
- Camera Calibration and 3D Reconstruction ↗OpenCV · Read 8 October 2026
Hand-eye calibration definition.
- ROS 2 tf2 coordinate frames ↗ROS 2 project · Read 8 October 2026
Official documentation source. Describes timestamped relationships between head, body, gripper and world coordinates. No claim that every humanoid uses ROS.
- ORB-SLAM3 Visual, Visual-Inertial and Multi-Map SLAM ↗Carlos Campos and colleagues, Universidad de Zaragoza · Read 8 October 2026
Accepted IEEE Transactions on Robotics paper, 2021. Used for visual odometry, map reuse and visual-inertial estimation. No claim that Atlas uses ORB-SLAM3.
- OctoMap A Probabilistic 3D Mapping Framework Based on Octrees ↗Armin Hornung, Kai M. Wurm, Maren Bennewitz, Cyrill Stachniss and Wolfram Burgard · Read 8 October 2026
Autonomous Robots, 2013. Free, occupied and unknown volume representation and updates from range measurements.
- Flipping the Script with Atlas ↗Boston Dynamics · Read 8 October 2026
Historical hydraulic Atlas parkour architecture. Manufacturer account of a 15 fps time-of-flight camera, plane segmentation, tracked obstacle geometry and an approximate course map.
Article history
Defined RGB-D at first use and added the documented sunlight limit of the indoor L515 depth sensor.
Report a correction