acceptodds
Under review as a conference paper at ICLR 2027

ProMac: Language-Guided World Action Modeling for Single-View Egocentric Navigation

Abstract

Egocentric visual navigation requires an agent to infer how to move through its surroundings toward a desired destination. We formulate this problem as predicting a feasible sequence of 2D waypoints from a single current egocentric image with a marked goal point. The goal specifies where to move, but reaching it requires reasoning about traversable space and intermediate motion while anticipating how the scene will evolve. This motivates representations that jointly capture semantic transition understanding and action-conditioned visual dynamics. We introduce ProMac, a world action model for egocentric navigation that uses an auxiliary chain-of-thought (CoT) reasoning reconstruction objective only at training time to guide world modeling. ProMac trains the visual token hidden states of a large multimodal model (LMM) backbone through two complementary pathways: (i) a frozen auxiliary language decoder provides supervision that compresses CoT reasoning information into the visual hidden states; (ii) a next-frame latent prediction head uses these “CoT-conditioned” visual patch hidden states to predict future visual token representations from a pretrained vision encoder. This language-guided world modeling objective grounds semantic transition reasoning in predicted visual dynamics, encouraging the trajectory policy to capture both how to move and how the scene will change under that movement. The reasoning and world modeling modules are used only during training. At inference, the vision-language model directly generates the 2D trajectory without explicit CoT generation or latent rollouts. On Ego4D Fut-Loc, ProMac reduces ADE, FDE, L2@First4, and DTW by 33.4%, 59.2%, 13.2%, and 49.1%, respectively, relative to trajectory-only Qwen3-VL-4B.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.