A robot that understands shape but not scale is not spatially intelligent. It is making an educated guess about whether the cup is within reach.
A new preprint introduces Metric-Aware Geometry Perception, or MAGP, as a plug-in framework for reconstructing scenes at real-world scale. The researchers argue that many existing geometry systems recover relative relationships while leaving the absolute scale arbitrary. The scene looks coherent, but object dimensions and distances can change across viewpoints and sensor configurations.
MAGP uses camera parameters and depth observations to anchor reconstruction to metric scale, then accepts different combinations and numbers of views. In experiments on ETH3D, MegaDepth, and ScanNet++, the authors report reducing absolute error from 2.01 meters to 0.07 meters while maintaining strong relative geometry accuracy.
When integrated into robot policies, the method improved performance across LIBERO, RoboTwin, and zero-shot LIBERO-Plus, with gains up to 6.26 percent on RoboTwin. Those are benchmark results reported by the authors in a preprint that has not completed peer review. They do not yet prove the same gains under messy lighting, sensor drift, deformable objects, or long deployments.
The direction is still exactly right. Embodied AI cannot live forever in normalized coordinates where everything is roughly the right shape. Manipulation depends on centimeters, collision margins, reach limits, and force applied at a real location. Better reasoning matters. Giving the robot a world that keeps the same size when the camera moves matters first.
LaunchPad positionRelative shape is enough for a pretty reconstruction. Physical action needs stable dimensions and distances, because a robot must know whether an object is seven centimeters away or seventy before it moves.
This report draws on the linked primary sources and reputable reporting. Company statements are treated as claims until independently demonstrated.
