Gae-WAM: Geometric Action Effect Modeling for World Action Models
Abstract
World action models (WAMs) use future visual prediction to support robot action learning. Complete future observations contain object motion, backgrounds, and textures, but not all of this information is directly relevant to action selection.How can action prediction emphasize task-relevant information in visual predictions?We propose Gae-WAM, a WAM architecture that connects visual and action predictions through action-induced geometric changes in target objects. Modeling these changes as geometric action effects brings object-focused information from visual predictions into action-side supervision without substantial additional inference overhead. We find that geometric action effect modeling is effective for robot manipulation. We compare Gae-WAM with other methods on multiple simulation benchmarks and Astribot S1 real-world tasks. Gae-WAM achieves higher aggregate success, with gains under camera-viewpoint and spatial-configuration changes. Shared-representation and training-objective ablations support the contributions of geometric references and cross-branch consistency to manipulation performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.