Learning Temporally-Aware Compact Gaussians for Feed-Forward 4D Reconstruction
Abstract
Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict per-pixel 3D Gaussians for each frame, leading to duplicated Gaussians and a view-dependent bias that hinder learning scene motion. We present Temporally-Aware Compact 4D Gaussian Splatting (T4GS), a feed-forward 4D reconstruction framework built upon a compact set of timestamp-conditioned learnable query tokens. Each token aggregates features across the full temporal context and decodes a 3D Gaussian whose position is modulated by the target timestamp, enabling globally coherent motion modeling. The resulting temporally coherent Gaussian representation serves as a motion-aware anchor for camera-controllable video generation and supports 4D feature lifting for point tracking and dynamic scene understanding. Across multiple dynamic benchmarks, T4GS achieves strong novel-view synthesis without input camera poses while using significantly fewer Gaussians and remaining robust to large temporal gaps.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.