Can Vision-Language Models Control Drones? Diagnosing Viewpoint-Aware Spatial Grounding
Abstract
Vision-language models (VLMs) are increasingly viewed as promising interfaces for aerial systems, where they could translate human goals and visual observations into drone commands. However, it remains unclear whether they can reliably map image-grounded spatial relations to directions defined in the aircraft movement frame. To address this gap, we introduce DroneBench, a diagnostic benchmark for one-step command prediction from an aerial image and a result-oriented task goal. DroneBench contains 1,350 paired images and tasks across four operational scenarios and two viewpoints, with each command decomposed into speed, aircraft-centered direction, and action for component-wise evaluation. Experiments with 15 VLMs show that models generally recognize task-relevant targets and choose plausible operational actions, yet remain substantially less reliable at direction grounding, especially for top-down views. We further evaluate five models on 50 matched forward-facing/top-down pairs collected during physical drone flights and observe the same failure pattern on onboard imagery. These results identify viewpoint-aware coordinate conversion as a central obstacle to reliable aerial command selection and establish DroneBench as a practical benchmark for advancing spatial reasoning in aerial VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.