G4D-Hand: Geometry-grounded 4D hand trajectory estimation from egocentric videos
Abstract
Egocentric 4D hand trajectory estimation aims to recover the 3D positions of hand keypoints over time in a consistent coordinate frame from egocentric video. Most existing methods follow an articulation-centric pipeline in which hands are detected and cropped, reconstructed from independent image patches, assigned wrist translation through a weak-perspective camera, and subsequently placed in the world by an external SLAM system. At no stage do these pipelines represent the hands in relation to the scene in which they move, even though this could help resolve hand depth and occlusion. We argue that scene geometry should shape the hand representation from the outset and introduce G4D-Hand, which grounds hand prediction in a geometric representation. It reconstructs both hands from shared full-frame features, without a hand detector or hand-specific image encoder, jointly recovering articulation and metric wrist translation in a common camera frame. Inter-hand and temporal supervision further encourage consistent relative positioning and coherent motion, and the same backbone estimates camera motion to lift predictions into the world frame without a separate SLAM stage. G4D-Hand runs at 30+FPS for camera-frame, matches state-of-the-art articulation accuracy, and substantially improves localization on egocentric benchmarks. On ARCTIC, a dataset with frequent hand occlusions, it reduces inter-hand relative wrist error by 50% and camera-frame hand trajectory error by 31%, against the strongest baseline on each metric. Code and checkpoints will be made public.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.