AnySpaceTime: Feed-Forward Spacetime Novel View Synthesis from Arbitrary Observations
Abstract
Novel view synthesis (NVS) has progressed significantly, but existing settings are typically tailored to specific observation patterns, such as static multi-view scenes, dynamic videos, or synchronized multi-camera captures. This paper introduces Space-Time NVS, a unified task that takes timestamped and posed images distributed across time and viewpoint as input and synthesizes images at a target viewpoint and a target time within the observed temporal interval, subsuming the above settings as special cases. We propose AnySpaceTime, an end-to-end feed-forward framework for this task that supports independently specified time and view queries in a single forward pass. Our key idea is to decouple spatial representation from temporal adaptation: we freeze a pretrained NVS encoder to preserve strong spatial priors and introduce a lightweight motion branch that predicts time-dependent updates in the latent space. The target time determines the latent update, while the target camera pose determines rendering, with static scenes handled as the shared-state special case. This design handles static multi-view inputs, dynamic videos, and asynchronous multi-camera observations without requiring a fixed camera layout, synchronized or contiguous frames, or supervision from depth, optical flow, or 3D motion. It generalizes to unseen scenes and irregular observations, and experiments across multiple datasets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.