Beyond Expert Trajectory Matching: Geometry and Risk-Aware Imitation Learning for Safer UAV Vision-Language Navigation
Abstract
UAV vision-language navigation (UAV-VLN) requires UAVs to follow natural-language instructions while avoiding obstacles using visual observations. A common approach to UAV-VLN is imitation learning, where navigation models are trained to mimic actions from expert demonstrations. However, vanilla imitation learning cannot distinguish multiple expert trajectories in terms of task-relevant qualities, and one particularly important quality for UAV navigation is safety. In reality, expert trajectories include ones that remain comfortably in open spaces, and ones that pass dangerously close to obstacles. Regardless of the difficulties in obstacle avoidance, vanilla imitation learning treats these trajectories as equally desirable supervision, which is not helpful for lowering collision rates. In order to address this, we adopt a simple but effective approach named GeoRisk-IL. Specifically, we introduce a collision-risk penalty on predicted actions using the distance between future UAV locations and obstacles. An essential aspect of our design is to also enhance the 3D awareness of the UAV-VLN model. This provides the spatial understanding needed for meaningful risk estimation. Interestingly, the collision-risk penalty becomes more effective when combined with geometric representation alignment, despite the latter providing little to no benefit on its own, highlighting a strong synergy between geometric awareness and risk-aware learning. On both synthetic and real-world datasets, the proposed method demonstrates significant improvements over previous methods, verifying its effectiveness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.