AerialWAM: Reconstructive Aerial World Action Model for Open-world Semantic-level Object-goal Search
Abstract
Open-world semantic-level object-goal search is an important yet challenging task for aerial agents. Existing approaches rely on spatio-semantic representations and multimodal large models to reason over historical observations, failing to maintain semantic-object alignment during long-horizon search in open-world environments. To solve this problem, this paper presents AerialWAM, the first reconstructive aerial world action model, which can autonomously search for semantic-level object-goals with semantic-object alignment. Specifically, AerialWAM first constructs a object-conditioned aerial world action model that couples semantic understanding and action generation using a Mixture-of-Transformers to achieve long-horizon prediction and action execution. We further design a semantic-object reconstructor that generates an goal-centric object image from current observations and the semantic description. During search, each new observation refines this object image and feeds it back to the world action model for the next action, forming a closed loop to achieve semantic-object alignment. We conducted extensive experiments on AerialDojo-200K. The result shows AerialWAM achieves average success rates 5.77× and 23.98× those of existing methods in in-distribution and out-of-distribution settings, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.