acceptodds
Under review as a conference paper at ICLR 2027

Throw D4RT into the Future: Motion World Models from 4D Reconstruction Prior

Abstract

Feedforward 4D reconstruction models learn a spatiotemporal prior of dynamic scenes from large collections of 4D-annotated video. Motion forecasters, in contrast, still rely on 3D trajectory annotations that are expensive to collect and limited in domain. We present FutureD4RT, a world model that forecasts in the feature space of a frozen 4D reconstruction model and is trained on raw video and timestamps alone. A tokenizer compresses each transition between consecutive scene states into one continuous delta token, a predictor forecasts future tokens, and the frozen decoder of the reconstruction model reads 3D positions from the predicted states. We build the same pipeline on two backbones, D4RT and Point4D. The delta token carries transition-specific content, since the correct token lowers reconstruction loss by 95.1% against copying the preceding state while tokens of other transitions do not. Forecasts beat copying the last state in feature space, and updating the random query vectors of the predictor from the waypoints of one supplied track lowers the average 3D error of the other points on PointMotionBench by 12 to 22% with D4RT and 6 to 20% with Point4D. These gains use an encoding of the full clip, as in training. From three observed frames alone, the forecast stays at the level of holding the last observed position, which we report as the main limitation. The results indicate that a frozen 4D reconstruction model can supply the state space of a motion world model without motion labels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.