Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
Abstract
Multimodal large language models (MLLMs) are increasingly used to make decisions for drones, but existing evaluations are mostly limited to simulation, to a single drone, and to targets that are visible from the start and do not move. We present DroneCATS, a benchmark in which the MLLM alone controls the drone through four actions, and evaluate ten MLLMs from frontier models to small open models. The benchmark varies target visibility, target motion, scene and simulator, and extends to real drones. Beyond a single drone, it adds a multi-drone suite, where one model drives multiple drones, and a commander + sub-agents suite, where a commander model directs one sub-agent for each of the drones. Across these settings, we find that frontier models achieve the highest success rates, but their latency causes them to lose moving targets, while small open models often do not declare arrival even when they reach the target. Guided by this failure analysis, we fine-tune a 2B model on drone action data synthesized from public RGB-D images; its success rate rises from 0% to 31% in simulation and outperforms the frontier models in real flight. With multiple drones, we find that every zero-shot model succeeds more often as the commander of one sub-agent per drone than when it drives all of them from one context, and that commanding success follows general VLM ability. On real drones, we observe that the results broadly follow those in simulation, while latency and hardware limits matter more. We will fully open-source the benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.