FlowEgoMotion: Scene-Conditioned Rectified Flow for Egocentric Full-Body Motion Estimation
Abstract
Estimating full-body motion from a head-mounted egocentric device is severely underconstrained: visual–inertial SLAM gives reliable egomotion estimation, but most of the wearer's body is never directly visible, and prior diffusion-based approaches incur high inference cost while underusing scene context. We present FlowEgoMotion, a scene-conditioned rectified-flow model that combines the SLAM head trajectory with two complementary scene pathways – time-aligned local geometry fused into the egomotion stream, and global tokens encoding room-scale structure – to condition a Transformer that predicts a motion vector field, while global root motion is recovered analytically from the trajectory. On Nymeria, a controlled comparison with the same backbone shows that scene conditioning reduces MPJPE from to mm, Wrist-PE from to mm, and floor penetration from to mm, and the gains carry over to ADT without fine-tuning. FlowEgoMotion also improves on the closest published scene-aware system, HMD, on the same benchmark, reducing MPJPE from to mm, Wrist-PE from to mm, and floor penetration from to mm without using the RGB features that HMD relies on. With only two flow steps, it needs about ms of GPU compute per second of generated motion and is more accurate than a matched 30-step diffusion model at lower cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.