acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Volumetric Reasoning for Aerial Vision-Language Navigation

Abstract

Aerial vision-language navigation (AVLN) requires unmanned aerial vehicles (UAVs) to follow natural-language instructions and reach specified destinations in large-scale, continuous 3D environments. Navigation-relevant cues are often distributed over long distances and observed at varying spatial scales, placing substantial demands on an agent's capacity for 3D spatial understanding and reasoning. This work introduces Hierarchical Volumetric Reasoning for AVLN, enabling the agent to reason over 3D scene structure from global layouts to fine-grained local regions. The reasoning is grounded in a Multi-Scale Voxel Map that organizes the observed environment into a coarse-to-fine 3D volume hierarchy and encodes geometric and semantic cues at each level. At each level, an instruction-conditioned decision token aggregates relevant voxel features, while its attention scores route subsequent reasoning to finer spatial regions. The final decision token is combined with navigation history to predict the next action. Experiments on AerialVLN-s show that our agent achieves 16.2% SR and 12.5% SDTW on seen environments, and 8.8% SR and 3.9% SDTW on unseen environments, with an inference latency of 0.09s per navigation step. Extensive ablation studies further validate the effectiveness of our core designs. The code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.