acceptodds
Under review as a conference paper at ICLR 2027

Can Vision-Language Models Control Drones? Diagnosing Viewpoint-Aware Spatial Grounding

Abstract

Vision-language models (VLMs) are increasingly viewed as promising interfaces for aerial systems, where they could translate human goals and visual observations into drone commands. However, it remains unclear whether they can reliably map image-grounded spatial relations to directions defined in the aircraft movement frame. To address this gap, we introduce DroneBench, a diagnostic benchmark for one-step command prediction from an aerial image and a result-oriented task goal. DroneBench contains 1,350 paired images and tasks across four operational scenarios and two viewpoints, with each command decomposed into speed, aircraft-centered direction, and action for component-wise evaluation. Experiments with 15 VLMs show that models generally recognize task-relevant targets and choose plausible operational actions, yet remain substantially less reliable at direction grounding, especially for top-down views. We further evaluate five models on 50 matched forward-facing/top-down pairs collected during physical drone flights and observe the same failure pattern on onboard imagery. These results identify viewpoint-aware coordinate conversion as a central obstacle to reliable aerial command selection and establish DroneBench as a practical benchmark for advancing spatial reasoning in aerial VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.