VerdictGS: Views as Noisy Annotators for Query-Time Open-Vocabulary Gaussian Splatting Segmentation
Abstract
Existing methods for open-vocabulary segmentation with 3D Gaussian Splatting integrate multi-view language features into 3D representations before any query. This adds aggregation, registration, and often grouping costs; moreover, query-independent aggregation mixes observations before their relevance to the query is known, predefined groups constrain the target's extent, and thresholds must be tuned before the query yet still ignore query-specific score distributions. We propose VerdictGS, a query-time inference framework that builds a 3D relevance field directly from individual observations' responses to a text query. Treating each view as a noisy annotator, VerdictGS aggregates per-view text–mask responses onto Gaussians and fits a per-query two-population binomial mixture to them with EM and without labels, yielding membership evidence whose decision boundary adapts to the query; this evidence is then fused across mask granularities. Because every decision is made at query time, VerdictGS builds no language representation before the query, lets each query weigh the original observations rather than their mixture, determines the target extent per Gaussian rather than within predefined groups, and needs no threshold tuning. Reducing this inference to sparse matrix operations and low-dimensional model fits, VerdictGS prepares scenes faster than training-free baselines and answers queries in comparable time, while achieving strong segmentation accuracy under each benchmark's evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.