acceptodds
Under review as a conference paper at ICLR 2027

One 4RROW, Every Target: for 4D Reconstruction and Recognition

Abstract

Feedforward models have made rapid progress on reconstructing 4D geometry and 4D tracking from dynamic video, and separately on incorporating semantics into static 3D reconstructions. However, no single model combines all three tasks. This paper introduces arrow, the first feedforward model to jointly perform 4D reconstruction, point tracking, and recognition from a single input video of arbitrary length, enabling extraction of semantic 4D instances of dynamic objects. Following D4RT, we use a powerful encoder coupled with an efficient point-wise decoder. We introduce object queries for instance recognition in such a way that they seamlessly operate alongside the point queries, and generalize the point query mechanism to videos of arbitrary length. To enable strong semantics we adapt DINOv3 for video and make it temporally aware, turning it into TimeDINO. On combined geometry and semantic prediction, arrow outperforms prior work on the static Scannet++ dataset while establishing a baseline on the dynamic Waymo dataset. Finally, we show our unified model can even outperform specialized video instance segmentation models in an open-vocabulary setting while also producing strong 4D reconstruction results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.