acceptodds
Under review as a conference paper at ICLR 2027

EgoH4RT: Bringing Egocentric Hands and Camera 4D Reconstruction Together

Abstract

Reconstructing hands from egocentric video requires preserving fine hand articulation while recovering camera motion in a consistent world coordinate system. This becomes more difficult over long videos, where errors in local estimates can accumulate. To address these challenges, we introduce EgoH4RT, a framework that combines the articulation priors of a pretrained hand model with the geometric understanding of a 3D foundation model. Hand queries connect these complementary representations, drawing on evidence across frames to recover detailed bimanual poses alongside metric camera motion. To carry these local predictions into a coherent reconstruction of the full video, we jointly optimize camera constraints spanning short and long time intervals, then refine world-space hand motion within the camera path. The model learns from approximately 2,600 hours of video across 14 egocentric sources through a training recipe that accommodates different joint definitions and incomplete annotations. Experiments show that EgoH4RT outperforms multi-stage reconstruction pipelines in world-space hand and camera reconstruction and methods built on pretrained video diffusion models in camera-space hand reconstruction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.