TTA4D: Object-Centric Free-Viewpoint Video Generation via One-Shot Test-Time Adaptation
Abstract
Generating free-viewpoint videos of a dynamic object from a single monocular video requires temporal coherence and view-consistent geometry, yet explicit 4D methods often distort geometry on in-the-wild inputs, and controllable video models drift from the input's appearance under large viewpoint changes. We present TTA4D, a one-shot test-time adaptation framework that converts a pre-trained image-to-video diffusion model into an object-specific free-viewpoint video generator in approximately 10–15 minutes on a single GPU. Our key observation is that the missing supervision, a moving object seen from a moving camera, factorizes into two readily available signals: a dynamic-object fixed-camera signal from the input video, which teaches the model to follow geometric conditioning as the object moves, and a static-object dynamic-camera signal rendered from a single-image 3D reconstruction, which teaches it to follow conditioning as the viewpoint changes. After LoRA adaptation on both, the model composes them along user-specified object-centric trajectories while preserving the input's appearance and motion. Because TTA4D adds only LoRA weights, without architectural changes or paired control data, it is directly applicable to newly released video foundation models. With this adaptation, a plain image-to-video model outperforms the depth-controllable VACE given the same geometry, and adapting VACE itself yields further improvement. On Consistent4D, TTA4D with the VACE backbone outperforms six baselines on all four metrics for both seen and unseen trajectories, and its outputs can directly supervise a 4D Gaussian Splatting reconstruction. On 22 real-world DAVIS videos, even its weakest configuration, a plain image-to-video model with coarse voxel geometry, receives a mean rating of in a blinded user study, compared with for the strongest baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.