acceptodds
Under review as a conference paper at ICLR 2027

EgoWAM: Geometry-Grounded World Action Models for Egocentric Interaction

Abstract

World models are typically learned in image-space state representations, such as pixels or 2D features. In egocentric hand-object interaction, however, such view-dependent states are heavily perturbed by arbitrary head motion, causing learned dynamics to entangle physical interactions with incidental viewpoint shifts. We present EgoWAM, a world action model that learns geometry-grounded dynamics in a 3D semantic field. This field represents the world by anchoring dense semantic features to scene geometry to reduce the influence of viewpoint changes. EgoWAM autoregressively predicts future 3D semantic fields and hand motions, with a diffusion-based action head that models uncertainty in future hand motions and provides a context-aware prior for sampling candidate actions during planning. We then score candidate actions using a spatial-semantic optimal transport objective, which establishes soft correspondences between predicted and goal states by aligning semantic features under a 3D transport cost. Experiments on both rigid and articulated egocentric hand-object interaction datasets show improved performance in world-state prediction, motion forecasting, and goal-directed planning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.