TTT4D: Mining Motion Cues in Test-Time Training for Scalable 4D Reconstruction
Abstract
Reconstructing dynamic 4D scenes from monocular videos remains challenging because moving objects violate the static scene assumptions of most 3D foundation models and interfere with camera and geometry estimation. Existing 4D methods often rely on fine tuning with limited dynamic data or external priors, while recent training-free approaches (e.g., Easi3R and VGGT4D) remain constrained by limited temporal scalability. In this work, we introduce TTT4D, a training-free framework that adapts the stateful 3D foundation model ZipMap to dynamic 4D reconstruction through test-time training. We observe that independent motion manifests as complementary inconsistencies in both ZipMap's internal representations and its online adaptation dynamics. Individual signals, however, are not motion specific and may also respond to static ambiguities such as newly observed regions, rich appearance, or uncertain geometry. Crucially, these ambiguities affect different internal signals differently, whereas independently moving regions exhibit consistent cross cue responses. Based on this observation, we extract complementary adaptation- and representation-level cues and fuse them to suppress ambiguous static responses and highlight coherent dynamic regions. The resulting masks guide a shallow attention routing strategy that selectively isolates dynamic features to reduce dynamic interference without retraining the pretrained model. TTT4D achieves strong performance in dynamic segmentation, camera pose estimation, and dense 4D reconstruction while retaining the efficiency and scalability of stateful feed-forward architectures, with speedups over VGGT-4D and over Easi3R on TUM-Dynamics / ADT benchmarks, respectively, using a single NVIDIA RTX PRO 6000 (96GB) GPU. Code will be released to contribute to the community.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.