acceptodds
Under review as a conference paper at ICLR 2027

NaviWAM: Coupled World–Action Prediction for Embodied Navigation

Abstract

Embodied navigation requires agents to translate task specifications and evolving visual observations into executable motion. Action labels supervise lowdimensional trajectories but do not directly constrain the richer visual changes along a route. We introduce NaviWAM, a coupled world–action prediction framework that complements trajectory supervision with patch-level next-frame prediction. A vision–language model integrates the task with long-horizon visual history to condition a shared dual-stream predictor. The predictor supports two separately invoked tasks: trajectory prediction (Policy) and next-frame feature prediction (NF). Shared parameters and cross-stream attention allow NF supervision to update components used for navigation. Labeled navigation sequences supervise both tasks, while additional video–instruction pairs provide NF targets without trajectory annotations. The same objective supports offline target-scene adaptation: target videos contribute NF supervision, while source trajectories retain Policy supervision through replay. At deployment, only the Policy task is sampled, without generating future visual observations. With separately trained vision-andlanguage and point-goal models, NaviWAM achieves competitive results across continuous navigation benchmarks. The combined Policy + NF training scheme increases RxR-CE success rate from 58.9% to 63.6%. On VLNVerse, target-scene NF adaptation improves fine- and coarse-grained success rates by 6.91 and 7.06 percentage points, respectively, without target-scene trajectory supervision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.