AheadFly: Think and Act Ahead for Multi-Stage Aerial Vision-and-Language Navigation
Abstract
Aerial vision-and-language navigation (AVLN) requires unmanned aerial vehicles (UAVs) to follow natural-language instructions in open 3D environments. Multi-stage navigation requires sequentially grounding multiple landmarks. Existing zero-shot approaches treat each stage as an independent grounding problem and primarily address what to look at in each stage. However, we argue that the dynamic nature of UAVs introduces additional challenges: aerial navigation must also resolve "where to look" and "where to look from" while maintaining grounding continuity across stages. In this paper, we propose AheadFly, a training-free modular framework that thinks and acts ahead for multi-stage aerial navigation. First, it "thinks ahead" by parsing instructions into explicit landmark queries and inter-stage transition directions. It then "acts ahead" by jointly adjusting the UAV's orientation and flight path under the guidance of these directions. Specifically, a pre-grounding view adjustment decides where to look by reorienting the UAV so the current landmark enters its field of view, and a lookahead path adjustment decides where to look from through multiple guiding forces, to mitigate cross-stage landmark occlusion and avoid obstacles along the path. To evaluate intermediate grounding and continuity, we introduce LandmarkFly, a multi-stage, long-range AVLN benchmark generated through a landmark-centric pipeline with explicit instruction-trajectory-landmark alignment, together with stage-level evaluation metrics. Experiments on both OpenFly-AirSim and LandmarkFly show consistent gains over the strongest baselines, improving Success Rate by 5.5 and 32.4 percentage points on the two benchmarks, respectively, and Grounding Point Accuracy by 7.6 percentage points on LandmarkFly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.