acceptodds
Under review as a conference paper at ICLR 2027

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

Abstract

Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations evolve under its predicted motion. Generative world-action models (WAMs) jointly predict future observations and actions, yet existing WAMs for continuous VLN do not condition joint future-view and action generation on geometry-aware representations inferred from the observation history. We present WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN. A frozen feed-forward geometry encoder and trainable 3D Scene-to-Token Adapter convert monocular RGB history into a fixed-length geometry-aware prefix that persistently conditions joint future-view and action generation. We train WNM-3D with supervised world-action learning, DAgger-style adaptation on policy-visited states, and Counterfactual DanceGRPO refinement. Across 3DGS-based GN-Bench and VLN-CE, WNM-3D outperforms the backbone-matched WNM-2D counterpart and strong VLN baselines on GN-Bench. We further demonstrate closed-loop deployment on a physical robot, while component and stage-wise ablations support the Adapter design and the complementary roles of DAgger-SFT and DanceGRPO.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.