PointSplatV2 : Query-Based Feed-Forward 4D Gaussian Reconstruction
Abstract
On-the-fly dynamic reconstruction is crucial for real-time spatial computing and immersive applications, requiring 4D representations that render stably from free viewpoints. Existing feed-forward methods rapidly reconstruct scenes by predicting Gaussian representations aligned with individual input observations, such as views or timesteps. This observation-wise discrete prediction prevents them from directly consolidating multi-view and temporal observations into a shared continuous, compact 4D representation, resulting in temporal flickering and redundant Gaussian primitives. Conversely, optimization-based methods achieve spatiotemporal consistency and compactness, but their costly per-scene optimization limits low-latency applications. To bridge this gap, we introduce PointSplatV2, a query-based feed-forward framework for reconstructing 4D Gaussians. Given multi-view temporal observations, PointSplatV2 first uses a multi-view image encoder to construct a latent field from the input frames. It then builds key-spatiotemporal queries that sparsely yet completely cover the spatial projection of the 4D scene at a key timestep. Each query then aggregates features from this latent field via cross-attention and directly decodes a 4D Gaussian representation. This simple yet effective formulation gathers temporal observations onto shared queries to avoid dense per-frame discrete prediction, and lets the continuous representation render stably across time. To extend this formulation to longer sequences, we introduce the Boundary-Aware Sequential Training (BAST) strategy, which enforces temporal coherence at sequence interfaces. Experiments show that PointSplatV2 outperforms the evaluated feed-forward 4D baseline in quality and compactness at low latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.