Hierarchical Reasoning-Guided Progressive Aerial Visual Grounding
Abstract
Complex aerial visual grounding requires not only understanding what a referring expression describes, but also coordinating how its different semantic components guide visual localization. This challenge is amplified in UAV imagery by small objects, dense same-category instances, and complex spatial layouts. We propose Hierarchical Reasoning-Guided Aerial Visual Grounding (HRVG), a unified framework that organizes semantic guidance throughout a progressive grounding process. The key idea is to align different semantic roles with corresponding visual computations, rather than treating semantic reasoning as an auxiliary language representation. HRVG derives hierarchical scene-, entity-, and relation-level semantic states from language reasoning and progressively integrates them into visual grounding, enabling contextual understanding, candidate grounding, and relational decision-making within a single connected framework. We further introduce RefScene-20K, a benchmark comprising 20,000 referring expressions over 8,550 UAV images across relational, ordinal, extremal, multi-target, and no-target tracks, with outputs uniformly represented as potentially empty sets of bounding boxes. Experiments on RefScene-20K demonstrate that HRVG outperforms existing grounding baselines. Further analyses, including semantic decomposition, reasoning-order, and stage-intervention studies, verify the effectiveness of organizing hierarchical semantics and aligning them with progressive visual computations. These results highlight semantic-role alignment as an effective modeling principle for complex aerial visual grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.