Mind the Orientation: Overcoming UprightBias in Feed-Forward Geometric Models
Abstract
We uncover a largely overlooked vulnerability in feed forward reconstruction models: sensitivity to image orientation. Across state-of-the-art architectures rotational perturbations of input images cause severe degradation in both camera pose estimation and point-map quality. We term this failure mode UprightBias and trace its origin to a near-universal upright orientation of training corpora, which causes models to implicitly assume canonically upright inputs at inference time. We find that patch-level embedding variance from a DINOv2 encoder provides a reliable zero-shot signal for identifying the canonical upright orientation. Leveraging this insight, we introduce ROVER (ROtation estimation via Variance-based Embedding Rectification), a training-free framework that estimates and rectifies image orientation at inference, leaving all downstream model weights unchanged. For evaluation, we introduce a Controlled Rotation Synthesis protocol that generates rotated views from upright images. We also curateRotMV (Rotational Multi-View), which, to the best of our knowledge, is the first benchmark of paired upright and real-world rotated multi-view sequences captured from fixed viewpoints. Experiments across the datasets demonstrate that ROVER consistently shows improvements, reducing pose estimation errors upto 80% and point-map estimation errors by over 58% on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.