AeroGround: Language Grounding in Large-Scale Aerial Gaussian Scenes
Abstract
Grounding natural-language queries in large-scale aerial 3D scenes requires distinguishing visually similar structures and recovering their complete 3D support. Large viewpoint changes and incomplete observations make both tasks challenging. We present AeroGround, a query-time framework that couples active view exploration with evidence-guided structural refinement. Active Query-guided View Exploration (AQVE) maintains competing cross-view target hypotheses and selects informative views to resolve ambiguity. It associates observations across views and accumulates visibility-aware positive and negative evidence to guide subsequent exploration. Evidence-Guided Structural Refinement (ESR) consolidates reliable foreground and background observations into Gaussian-level target posteriors and combines them with local geometric consistency to recover missing support while suppressing foreground leakage. AeroGround operates directly on standard 3D Gaussian Splatting without scene-specific semantic feature optimization or storage. We also introduce AeroRef3D, a benchmark covering spatial–attribute, commonsense, functional, and hypothetical queries with zero-to-many targets across six aerial scenes. Experiments show that AeroGround substantially outperforms the evaluated baselines on AeroRef3D and achieves competitive results on two aerial open-vocabulary segmentation benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.