One 4RROW, Every Target: for 4D Reconstruction and Recognition
Abstract
Feedforward models have made rapid progress on reconstructing 4D geometry and 4D tracking from dynamic video, and separately on incorporating semantics into static 3D reconstructions. However, no single model combines all three tasks. This paper introduces arrow, the first feedforward model to jointly perform 4D reconstruction, point tracking, and recognition from a single input video of arbitrary length, enabling extraction of semantic 4D instances of dynamic objects. Following D4RT, we use a powerful encoder coupled with an efficient point-wise decoder. We introduce object queries for instance recognition in such a way that they seamlessly operate alongside the point queries, and generalize the point query mechanism to videos of arbitrary length. To enable strong semantics we adapt DINOv3 for video and make it temporally aware, turning it into TimeDINO. On combined geometry and semantic prediction, arrow outperforms prior work on the static Scannet++ dataset while establishing a baseline on the dynamic Waymo dataset. Finally, we show our unified model can even outperform specialized video instance segmentation models in an open-vocabulary setting while also producing strong 4D reconstruction results.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.