Trajectory-Anchored Compact Gaussians for Feed-Forward 4D Reconstruction
Abstract
Dynamic scene reconstruction from monocular video requires a representation that supports both novel-view rendering and correspondence through time. Feed-forward methods learn reconstruction priors across scenes, enabling inference on unseen videos without costly per-scene optimization. However, current feed-forward methods rely heavily on dense pixel-aligned architecture which produces redundant primitives. On the other hand, recent compact query-based methods can struggle to maintain accurate temporal correspondence. We present a feed-forward framework for compact, structured 4D Gaussian reconstruction without input camera poses. Our central idea is to use dense correspondence to organize Gaussians around a sparse set of explicit motion trajectories. We first predict dense B-spline trajectories with a geometry-anchored parameterization that preserves each trajectory's source point position. Grouping trajectories by spatial proximity, motion, and appearance yields representative anchors. Given learnable tokens paired with the anchor trajectories, a query-based decoder aggregates image evidence to generate child Gaussians, whose local offsets capture geometry and motion beyond each anchor. This hierarchy supports rendering and dense 3D correspondence with a compact Gaussian representation whose size is independent of the number of input pixels. Across the evaluated benchmarks, our full model outperforms the evaluated pixel-aligned baselines in most rendering metrics while using significantly fewer Gaussians. The same compact representation outperforms the evaluated query-based baselines across all reported rendering and Gaussian-tracking metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.