acceptodds
Under review as a conference paper at ICLR 2027

GeoMVT: Rethinking Multi-View Point Tracking through Geometric Reasoning

Abstract

We introduce GeoMVT, a depth-free multi-view point tracker for accurate 2D tracking and reliable 3D motion recovery. Our key idea is to leverage calibrated 3D camera geometry to bridge multi-view 2D tracking and accurate 3D motion recovery by guiding cross-view interactions during trajectory estimation. To this end, we propose a geometry-guided cross-view reasoning framework with two complementary stages: correlation aggregation and point-state updates. We first design epipolar-guided geometric transform attention (EpiGTA) to aggregate local correlation features through relative geometric alignment and adaptive epipolar guidance. Building on these geometry-enhanced features, a geometry-guided transformer combines spatio-temporal context with geometrically aligned cross-view interactions. Together, the two stages form a recurrent loop, where geometry-enhanced correlations guide refinement and updated trajectories determine subsequent correlation sampling and geometric compatibility. We construct a synthetic multi-view dataset of 12K sequences using Kubric for training. Experiments on four standard benchmarks demonstrate substantial improvements over existing multi-view 2D trackers and effective use of additional views across different camera configurations. Moreover, simple triangulation yields competitive 3D tracking, surpassing dedicated 3D trackers on multiple real-world benchmarks without depth input.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.