Look Before You Localize: Active Target Localization for Aerial Vision-and-Language Navigation
Abstract
In language-goal aerial navigation, a UAV progressively acquires observations to localize a target described in a language instruction by its visual attributes and relations to surrounding landmarks. A target may already be visible while the relational, ordinal, or landmark context needed to distinguish it from plausible alternatives remains outside the observed region. This creates an important decision during navigation: whether to localize from the current visual context or continue observing the environment. We propose a closed-loop active target localization framework that makes this decision explicit through two complementary agents. The Selective Target Localization Agent either predicts a target location or defers localization based on the accumulated visual context, while providing a rationale associated with the deferral. The Rationale-Guided Exploration Agent conditions on this rationale and the shared spatial context to select the next exploratory movement. New observations update a world-aligned spatial memory, allowing localization to be reassessed using the retained visual context. Experiments on CityNav show consistent improvements across seen and unseen splits. On the original Test Unseen split, our method achieves 42.57% SR and 30.95% SPL, exceeding the best compared automated SR and SPL by 16.67% and 11.32%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.