Grow Where Views Agree: Efficient Sparse-View 4D Reconstruction with GI-ACE
Abstract
With only four cameras, a dynamic Gaussian-splatting reconstruction sees each surface from few directions, yet densification still adds primitives wherever the pooled image-space gradient is large. We argue that this signal conflates two quantities: how strongly a primitive asks to change, and how many independent viewpoints make the request. A residual that only one camera sees can be removed by fitting that camera, so it is weak evidence that new scene structure is needed. We introduce GI-ACE (Geometry-Informed Asynchronous Camera Evidence), an admission rule for 4D Gaussian densification that uses the per-camera gradient statistics ordinary training already produces. A primitive grows if its demand survives removing its strongest camera, or, for a large primitive, if two cameras jointly request a split from sufficiently different directions, measured at the times each camera observed it. Both tests reuse the native threshold and add no learned component, point budget, or rendering pass. On the six-scene Neural3DV four-camera benchmark, GI-ACE exceeds, to our knowledge, all published results, improving mean PSNR from 22.29 to 23.04 dB over the strongest baseline, and in our reproductions it uses 2.9–3.7 fewer Gaussians while training 1.6–2.0 faster than the same backbone under native densification. On six Panoptic Sports sequences it improves mean PSNR by 0.45 dB with an order of magnitude fewer Gaussians and 5.4 faster training. Controlled studies confirm the design: the population stays stable under biased camera sampling, temporally staggered capture further improves the same rule, and the geometric route activates precisely where complementary views exist.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.