acceptodds
Under review as a conference paper at ICLR 2027

AerialWAM: Reconstructive Aerial World Action Model for Open-world Semantic-level Object-goal Search

Abstract

Open-world semantic-level object-goal search is an important yet challenging task for aerial agents. Existing approaches rely on spatio-semantic representations and multimodal large models to reason over historical observations, failing to maintain semantic-object alignment during long-horizon search in open-world environments. To solve this problem, this paper presents AerialWAM, the first reconstructive aerial world action model, which can autonomously search for semantic-level object-goals with semantic-object alignment. Specifically, AerialWAM first constructs a object-conditioned aerial world action model that couples semantic understanding and action generation using a Mixture-of-Transformers to achieve long-horizon prediction and action execution. We further design a semantic-object reconstructor that generates an goal-centric object image from current observations and the semantic description. During search, each new observation refines this object image and feeds it back to the world action model for the next action, forming a closed loop to achieve semantic-object alignment. We conducted extensive experiments on AerialDojo-200K. The result shows AerialWAM achieves average success rates 5.77× and 23.98× those of existing methods in in-distribution and out-of-distribution settings, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.