Flow4RT: Seeing Motion through Geometry
Abstract
Reconstructing and tracking dynamic scenes requires choosing a reference against which motion is measured: the background and independently moving objects support different rigid alignments across views. Directly regressing global geometry or camera poses makes this choice implicit and difficult to change after prediction. We present Flow4RT, a feedforward model that extends Flow4R's local geometry and relative motion formulation from image pairs to joint multi-view inference. The model aggregates information across all input views and predicts motion in both directions between each view and an anchor, supporting reconstruction and tracking within one formulation. Our key insight is that motion is displaced geometry. We predict local geometry for each view, then condition its features on a target view to predict where the same points move in viewpoint and time. A single dense head decodes both local and displaced point maps with their respective confidence and rigidity weights, avoiding duplicated geometry predictions and separate output heads. A closed-form, rigidity-weighted alignment recovers the transformations needed for reconstruction and tracking without a pose regression head. Adjusting the alignment weights selects a different rigid reference object while reusing the same dense predictions. On WorldTrack, Flow4RT achieves the best average tracking performance, improving APD by 2.4 points and reducing EPE by 17% over the strongest prior result for each metric, while ranking second in reconstruction on both metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.