Geometry-Aligned Rewards For Aerial Vision-And-Language Navigation
Abstract
Aerial vision-and-language navigation (AVLN) requires an agent to localize a language-described goal on a city-scale aerial map by grounding visual landmarks and reasoning over their spatial relations. Recent vision-language model (VLM)- based methods formulate goal localization as coordinate generation and optimize the policy with reinforcement learning. Their composite rewards typically cal- ibrate goal accuracy with globally fixed distance scales, limiting discrimination beyond a cutoff and ignoring variation in landmark extent. To address this is- sue, we propose GS-Nav, a geometry-aligned learning framework for AVLN. Its core is a Scale-Adaptive Gaussian Reward, which converts the distance to the target into a dense preference signal and adjusts its bandwidth according to the spatial extent of the referred landmark. In addition, we identify a simple but ef- fective resolution-decoupling strategy: training with high-resolution maps while performing inference at a lower resolution. On the CityNav benchmark, GS-Nav obtains the best reported SR, OSR, SPL, and mean NE point estimates among the listed learned methods on validation unseen. Meanwhile, resolution-decoupled inference uses approximately 3.1×fewer visual tokens and reduces median query latency by about 4.2×. Controlled reward comparisons support geometry-aligned supervision within CityNav; resolution controls reproduce the inference trend at two scales of the Qwen2.5-VL family. Code and models will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.