AlignFly: Landmark-Alignment Reinforcement Fine-Tuning for UAV Vision-Language Navigation
Abstract
Vision-language navigation for unmanned aerial vehicles (UAV-VLN) requires agents to follow natural language instructions using rapidly changing visual observations along three-dimensional trajectories. Such instructions usually describe routes via ordered visual landmarks, which provide stage-wise feedback for navigation. Existing UAV-VLN research rarely explores the correspondence between landmarks and navigation stages. In this paper, we propose AlignFly, a landmark-alignment reinforcement fine-tuning (RFT) method for optimizing navigation actions. Starting from a vision-language model (VLM) initialized through supervised fine-tuning (SFT) on expert demonstrations, AlignFly samples candidate action chunks from the same expert state and executes them independently in a simulator. A VLM-as-Judge is introduced to compare the visual evidence of the target landmark in post-action observations of candidates and construct rewards for group relative policy optimization (GRPO). To provide the required target-landmark information, we develop an automated annotation pipeline that extracts ordered landmark phrases from instructions, scores their visual evidence in expert trajectory observations, and monotonically aligns them with trajectory stages through dynamic programming. We apply this pipeline to a scene dataset of OpenFly and use the resulting annotations for RFT. Experiments on the dataset demonstrate consistent improvements on the seen split and a held-out scene, while ablations support the contribution of landmark-alignment feedback to UAV navigation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.