SIMPLIFYING COMPLEXITY: STAGE ADAPTIVE EVIDENCE ROUTING FOR UAV VLN
Abstract
UAV vision-language navigation requires long-range search, target identification, and precise terminal control. However, existing methods typically use the same visual-language inputs throughout navigation, failing to account for the change in evidence requirements before and after target discovery. Before the target appears, navigation primarily relies on scene context and directional hints for exploration; once the target is reliably identified, the decision process should shift toward target-centered local control. To address this, we propose a stage-adaptive evidence-routing framework that divides navigation into exploration before target discovery and target-centered navigation afterward. The latter is further divided into approach, regrounding, and reacquisition. The framework combines hint-conditioned target identification with lightweight target memory to determine the current stage from the relationship between current and previously accepted target evidence. It then routes the corresponding visual and textual information to a shared action policy. On 3,005 TravelUAV test episodes, our method achieves success rates of 58.41%, 66.29%, and 40.71% on the Seen, Unseen Object, and Unseen Map splits, improving over AerialVLA by 10.45, 9.69, and 3.13 percentage points, respectively. These results show that adapting decision inputs to target discovery and changes in target evidence improves successful termination across diverse navigation conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.