: Towards Efficient Closed-Loop Aerial Vision-Language Navigation via Visual Compression
Abstract
Aerial vision-language-action (VLA) models incur substantial computational costs from visual feature extraction and downstream processing of dense visual token sequences, limiting navigation decision rates on resource-constrained unmanned aerial vehicles (UAVs). Existing approaches reduce these costs through encoder compression or token pruning, but optimizing either stage alone leaves the other source of computation largely unchanged. Moreover, token selection based solely on current-frame information can overlook cues needed for subsequent navigation, as flight actions change the viewpoint and available observations. We introduce Aerial Closed-Loop Navigation via Visual Compression (), a framework that addresses both computational costs while using navigation context to guide visual information retention. Complementary-teacher distillation transfers spatial and semantic features from two teacher encoders into a single lightweight student, which replaces both teachers during inference. A state-adaptive token allocator then uses navigation context and recent control history to select visual tokens, adjusts the token budget according to observation uncertainty, and applies role-aware retention constraints. The VLA backbone processes only the selected tokens, reducing downstream multimodal computation. We evaluate on the Seen, Unseen Object, and Unseen Map splits of TravelUAV. Compared with the dense VLA baseline, reduces FLOPs by 40.54% and mean end-to-end latency by 27.72% while also improving navigation performance. On the Seen split, it achieves a success rate (SR) of 51.48% and success weighted by path length (SPL) of 41.02%, representing improvements of 3.95% and 3.18%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.