acceptodds
Under review as a conference paper at ICLR 2027

AlignFly: Landmark-Alignment Reinforcement Fine-Tuning for UAV Vision-Language Navigation

Abstract

Vision-language navigation for unmanned aerial vehicles (UAV-VLN) requires agents to follow natural language instructions using rapidly changing visual observations along three-dimensional trajectories. Such instructions usually describe routes via ordered visual landmarks, which provide stage-wise feedback for navigation. Existing UAV-VLN research rarely explores the correspondence between landmarks and navigation stages. In this paper, we propose AlignFly, a landmark-alignment reinforcement fine-tuning (RFT) method for optimizing navigation actions. Starting from a vision-language model (VLM) initialized through supervised fine-tuning (SFT) on expert demonstrations, AlignFly samples candidate action chunks from the same expert state and executes them independently in a simulator. A VLM-as-Judge is introduced to compare the visual evidence of the target landmark in post-action observations of candidates and construct rewards for group relative policy optimization (GRPO). To provide the required target-landmark information, we develop an automated annotation pipeline that extracts ordered landmark phrases from instructions, scores their visual evidence in expert trajectory observations, and monotonically aligns them with trajectory stages through dynamic programming. We apply this pipeline to a scene dataset of OpenFly and use the resulting annotations for RFT. Experiments on the dataset demonstrate consistent improvements on the seen split and a held-out scene, while ablations support the contribution of landmark-alignment feedback to UAV navigation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.