Chronofy: Pose-Aware Generative 4D Reconstruction from Videos
Abstract
Recent generative reconstruction methods effectively reconstruct static 3D content but often exhibit pronounced inter-frame jitter when applied to dynamic content. We present Chronofy, a pose-aware 4D generative reconstruction framework that jointly predicts complete dynamic objects and unified camera-space poses from single-view videos. Built upon SAM3D, Chronofy transfers its powerful generative reconstruction priors to the 4D domain through three key designs: (i) a pose-aware 4D representation that encodes frame-dependent motion in spatiotemporal structured latents and uses a unified pose for sequence-level placement; (ii) a generative reconstruction model that utilizes a spatiotemporal mixture of transformers for temporal modeling and frame-wise correspondence between shape and pose; and (iii) a dynamic Gaussian decoder that combines temporal latent interactions with time-varying Gaussian attributes to reduce inter-frame flickering. To preserve occlusion robustness with limited 4D training data, we use temporal masking augmentation, while layout test-time optimization refines multi-object spatial alignment. Experiments on Motion-80, Kubric-4D, and our simulated real-world Chrono-64 datasets demonstrate that Chronofy substantially improves spatiotemporal consistency in both 4D object generation and 4D scene reconstruction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.