acceptodds
Under review as a conference paper at ICLR 2027

EgoAct4D: Estimating Bimanual 4D Hand Motion from Egocentric Videos

Abstract

Turning egocentric video into action supervision requires accurate and efficient hand motion estimation. Multi-camera capture and wearable sensing provide accurate motion supervision but are costly to scale. Monocular reconstruction could enable large-scale supervision from abundant Internet videos, but existing pipelines require multiple stages and incur motion errors of several centimeters. To address these limitations, we present EgoAct4D, a simpler pipeline for more accurate bimanual 4D hand motion estimation from monocular egocentric RGB videos. Its single-stage hand estimator predicts bimanual states from short clips, including occluded and out-of-view poses, while a separate estimator directly predicts metric camera poses. The pipeline reuses hand predictions and stitches overlapping camera trajectories to express hand motion in each window's first-frame camera coordinates and pair it with the starting observation. Since existing alignment-based evaluations can mask position and motion errors, we measure camera-space joint error (Absolute MPJPE) and hand-trajectory error (Action MPJPE) in sliding windows without ground-truth alignment. EgoAct4D achieves state-of-the-art results on both metrics on the HOT3D benchmark. We then build the EgoAct4D dataset and demonstrate full-pipeline data scaling beyond HOT3D's training scale. On its test set, EgoAct4D achieves 7.95 mm Absolute MPJPE and 10.98 mm Action MPJPE at an end-to-end throughput of 49.3 FPS. This accuracy and efficiency can support reusable motion supervision for vision-language-action pretraining and human-to-robot motion retargeting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.