Allocentric Memory, Egocentric Acting: Dual-Frame based VLA for City-scale Aerial Navigation
Abstract
Language-goal aerial navigation requires a UAV to reach a destination described in natural language using observations gathered along its flight. A central challenge is that observations and motion commands depend on the UAV's changing pose, whereas environmental landmarks provide fixed spatial references for long-range reasoning. To reconcile these reference frames, we propose a vision-language-action model for aerial city navigation, named CityNav-VLA. Its dual-frame design retains accumulated observations in allocentric spatial memory, with metric coordinates anchoring each map token to a fixed physical location. A navigation query combines instruction-conditioned perception with the UAV's state and motion history to retrieve relevant evidence from this memory and predict a waypoint. The waypoint remains in the map frame until execution, when a controller converts it into body-frame motion commands. Early goal commitment then limits the influence of later prediction drift by fixing the execution waypoint to the median of initial estimates. On CityNav Test Unseen, CityNav-VLA reduces navigation error by 30.6 m and increases success rate by 4.6 percentage points relative to the best competing result for each metric. Cumulative ablations show that navigation-map features provide the largest gains and that metric coordinates further improve endpoint accuracy. With a 4B-scale backbone, CityNav-VLA averages 0.68 s of model inference per decision on single GPU, reducing at least 47.8% GPU memory and 84.0% decision latency with the comparable LLM-based methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.