GALE: Learning Where to Distill for Sparse-View 3D Gaussian Segmentation
Abstract
Sparse-view 3D Gaussian segmentation aims to recover scene geometry and consistent object identities from only a few images. Existing methods use geometry and segmentation foundation models for this task, but do not fully resolve which Gaussians each 2D mask should supervise. This ambiguity arises because Gaussians on different surfaces can project into the same mask. In this paper, we present GALE (Geometry-Aware Lifting of Evidence), a lightweight framework that learns where to distill through geometry-conditioned sparse attention. Built on a VGGT-initialized reconstruction, GALE combines predicted depth and confidence with Gaussian visibility to automatically select and weight the 3D support of each SAM2 mask. These uncertainty-aware assignments focus supervision on compatible Gaussians, guiding knowledge transfer in both scene-specific optimization and cross-scene prediction. Experiments show that GALE outperforms prior distillation and grouping methods on views with held-out annotations. When generalizing to unseen scenes, it surpasses other methods on most metrics using only 1.51M task-specific trainable parameters, approximately 1/40 of the state-of-the-art method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.