GeLoc3r: Enhancing Relative Camera Pose Regression with Geometric Consistency Regularization
Abstract
ReLoc3R achieves competitive performance with fast 33ms inference and state-of-the-art regression accuracy, yet our analysis reveals subtle geometric inconsistencies in its internal representations that prevent reaching the precision ceiling of correspondence-based methods like MASt3R (which requires 294ms per pair). In this work, we present GeLoc3r, a novel approach to relative camera pose estimation that enhances pose regression methods through Geometric Consistency Regularization (GCR). GeLoc3r overcomes the speed-accuracy dilemma by training regression networks to produce geometrically consistent poses without inference-time geometric computation. During training, GeLoc3r leverages ground-truth depth to generate dense 3D-2D correspondences, weights them using a FusionTransformer that learns correspondence importance, and computes geometrically-consistent poses via weighted RANSAC. This creates a consistency loss that transfers geometric knowledge into the regression network. Unlike the FAR model, which requires both regression and geometric solving at inference, GeLoc3r only uses the enhanced regression head at test time, maintaining ReLoc3R's fast speed and approaching MASt3R's high accuracy. Against a vanilla ReLoc3R trained on the same data, GeLoc3r improves AUC@ from 34.20% to 41.42% on CO3Dv2 (21% relative), from 63.54% to 68.66% on RealEstate10K, and from 39.86% to 50.45% on MegaDepth1500. By teaching geometric consistency during training rather than enforcing it at inference, GeLoc3r offers a new perspective on how neural networks learn camera geometry, achieving both the speed of regression and the geometric understanding of correspondence methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.