Geometry Meets Structure: Learning Robust Monocular Visual Odometry with Scale-Preserving Constraints
Abstract
Self-supervised monocular visual odometry learns motion from appearance but lacks explicit geometric consistency at inference time. Two-view geometry provides complementary motion constraints, yet its translation is scale-ambiguous and its estimates can become unreliable under weak parallax or noisy correspondences. We investigate how learned motion predictions and explicit geometry can be selectively combined, and present GeoStruct-VO. Its Structure-Aware Feature Enhancement (SAFE) module augments shallow appearance features with spatially distributed responses for pose estimation. At inference time, Scale-Preserving Geometric Refinement (SPGR) assesses geometric reliability based on correspondence support, cheirality, epipolar consistency, and agreement with the learned prediction. Reliable estimates refine rotation and translation direction while retaining the network-predicted translation magnitude; otherwise, the learned pose is preserved. A Multi-frame Rotation Graph (MRG) further incorporates quality-filtered long-gap rotational constraints to reduce accumulated drift. Experiments on KITTI and cross-domain evaluations on vKITTI2 show that GeoStruct-VO improves translational and rotational accuracy over the evaluated learning-based methods. Component analyses further demonstrate the complementary effects of feature enhancement, selective geometric correction, and multi-frame rotation refinement. The real-time configuration processes each frame in 20.9 ms, showing that explicit geometry can improve learned odometry without online parameter adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.