DepthUMI: Active Depth for Universal Manipulation Interfaces
Abstract
Universal Manipulation Interfaces (UMIs) scale robot learning by allowing people to collect demonstrations beyond the robot workspace. Their portability, however, shifts the bottleneck to sensing: a wrist-contained device must recover precise motion despite occlusions and rapid viewpoint changes, while preserving the local 3D geometry needed by the policy. We present DepthUMI, a handheld interface in which the same learned active-depth source supports both motion recovery and policy observation. Rather than treating these as separate sensing problems, DepthUMI combines wide-field stereo, high-rate inertial sensing, and learned active depth in one calibrated assembly. Wide-field visual–inertial reconstruction estimates the demonstrated trajectory, and a calibrated transfer maps camera motion into gripper action supervision. In parallel, learned near-field geometry constrains reconstruction and supplies metric point-cloud observations for policy learning. A simulation study guides the sensing configuration, followed by real-world evaluation of the resulting system. On an instrumented validation rig evaluated against an external Vive reference, DepthUMI achieves an overall mean wrist-camera translation error below 3 mm without using that reference as estimator input. Across three manipulation tasks spanning articulated pulling, peg insertion, and constrained block extraction, learned geometry matches the strong baseline on the coarse pulling task and yields its clearest gains when local geometric precision is decisive. Together, these results show that a single wrist-contained sensing system can produce accurate action supervision and geometry-aware policy inputs. Project page: https://anonymous-submission-x.github.io/DepthUMI/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.