AeroEVA: Unified World Modeling and VLM Policy Learning for Embodied Aerial Search
Abstract
Efficient aerial search for human-specified goals requires anticipating what each move will reveal. Content next to the current view is often predictable even when distant regions are not, yet existing agents remain limited in turning such local foresight into better actions. We introduce AeroEVA (Envision, Verify, and Act), a vision-language model (VLM) agent that uses transition prediction in two roles: as a training signal, and as evidence checked before each move. To envision, its vision tower learns to predict features of the nearest patches in the neighboring region under each action, and textual transition targets supervise the decoder alongside actions. To verify, AeroEVA reuses this predictor to score each candidate move against an aerial-image goal. To act, the policy judges when these one-step scores are relevant and combines them with observations and history. Imitation learning alone surpasses prior search agents, and multi-turn reinforcement learning widens the margin to nearly 20 points in success rate. Without retraining, AeroEVA also leads prior methods in unseen regions, with ground-photo and text goals, under disaster-induced appearance change, and on finer grids. We further introduce AeroTask, 696 search tasks across 13 application scenarios whose language goals state task requirements rather than target appearance. Code, model weights, and the AeroTask dataset will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.