acceptodds
Under review as a conference paper at ICLR 2027

TracksTo4DGS: Feed-Forward 4D Gaussian Splatting from Dense Tracks

Abstract

Rendering dynamic scenes from monocular video has attracted growing interest, driven by applications in AR/VR, egocentric perception, and embodied AI. The task is challenging, as it requires jointly recovering accurate geometry, motion, and appearance from a single moving viewpoint. Recently, pretrained foundation models for tracking and reconstruction have made remarkable progress on the first two, providing robust 3D geometry and dense correspondences across diverse environments. However, their outputs are point clouds, which lack the representation needed to render high-quality video. We present TRACKSTO4DGS, an architecture that closes this gap by decoupling scene geometry and motion from appearance. Rather than learning temporal dynamics from scratch through an ambiguous rendering loss, TRACKSTO4DGS leverages frozen foundation models to explicitly define the 3D trajectory of each Gaussian, while a dedicated Attribute Network learns only the remaining properties required for high-fidelity rendering. To make dense tracking data computationally tractable, a sparse bottleneck reasons over a compact set of representative trajectories before decoding attributes for the full dense set. By inheriting the robust geometry of these foundation models, TRACKSTO4DGS achieves strong zero-shot generalization. Trained exclusively with a rendering loss on fixed-length egocentric video, a single model improves average PSNR by 2.0 dB over the strongest baseline across six diverse benchmarks, spanning egocentric, handheld, and object-centric captures, and seamlessly generalizes to unseen video lengths without per-dataset tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.