FL4N: Feed-forward Latent Model for 4D Novel View Synthesis
Abstract
4D Novel View Synthesis (NVS) from monocular videos aims to render dynamic scenes across viewpoints and timestamps, requiring jointly inferring 3D structure and temporal dynamics. Existing feed-forward methods rely on explicit 3D Gaussian representations, which are prone to temporal inconsistency and severe rendering artifacts. We present FL4N, the first feed-forward latent 4D NVS model. We bypass explicit 3D primitives by weaving a geometrically grounded scene prior into a unified latent representation across space and time. To overcome dynamic training data scarcity, we learn a temporal residual on top of a frozen, pretrained static latent backbone. A lightweight dynamic encoder learns to isolate true scene motion from camera parallax. Trained purely on image synthesis with binary mask guidance, FL4N achieves state-of-the-art rendering quality for 4D NVS in realtime. Our proposed latent representation successfully learns a strong 4D prior, generalizes across novel views and timestamps, and cleanly decouples view and time in the latent space. Codes will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.