acceptodds
Under review as a conference paper at ICLR 2027

Look Before You Localize: Active Target Localization for Aerial Vision-and-Language Navigation

Abstract

In language-goal aerial navigation, a UAV progressively acquires observations to localize a target described in a language instruction by its visual attributes and relations to surrounding landmarks. A target may already be visible while the relational, ordinal, or landmark context needed to distinguish it from plausible alternatives remains outside the observed region. This creates an important decision during navigation: whether to localize from the current visual context or continue observing the environment. We propose a closed-loop active target localization framework that makes this decision explicit through two complementary agents. The Selective Target Localization Agent either predicts a target location or defers localization based on the accumulated visual context, while providing a rationale associated with the deferral. The Rationale-Guided Exploration Agent conditions on this rationale and the shared spatial context to select the next exploratory movement. New observations update a world-aligned spatial memory, allowing localization to be reassessed using the retained visual context. Experiments on CityNav show consistent improvements across seen and unseen splits. On the original Test Unseen split, our method achieves 42.57% SR and 30.95% SPL, exceeding the best compared automated SR and SPL by 16.67% and 11.32%, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.