AnchorSplat: Feed-Forward 3D Referring Segmentation in Gaussian Scenes
Abstract
Grounding a referring expression in sparse, uncalibrated views requires both semantic discrimination and 3D consistency. Feed-forward reconstruction models supply cameras, depth, and renderable Gaussians, yet substituting them for the backbone of a multiview referring model consistently lowers accuracy. We trace this failure to two mechanisms. Reconstruction pretraining progressively attenuates the appearance cues that referring expressions rely on, and pixel-aligned Gaussian rendering allows each primitive to copy the label of the pixel that created it. We present AnchorSplat, which predicts view masks and a referred Gaussian scene in a single forward pass. A frozen semantic encoder anchors the referring state, and zero-initialized gates admit reconstruction features while preserving its initial referring function. A frame-invariant correspondence graph transports language-conditioned states across views, and an opacity-gated readout labels each Gaussian under owner-excluded rendering, which evaluates a primitive only in views that did not create it. On the 9,508 test samples of MVRefer, AnchorSplat improves global mIoU from 39.9 to 41.0 and [email protected] from 41.5 to 44.6 over MVGGT, whereas direct substitutions with VGGT-Ω, DA3, and AMB3R reach at most 32.8 global mIoU.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.