acceptodds
Under review as a conference paper at ICLR 2027

NaviVLN:Learning Task-Centric Hierarchical Representations via Embodied Alignment for Dual-System Vision-Language Navigation

Abstract

Vision-Language Navigation (VLN) integrates spatial perception, instruction understanding, and sequential decision-making. Foundation-model-based policies benefit from multimodal supervision, yet insufficient task relevance and cross-modal correspondence can cause task-agnostic intra-modal redundancy and cross-modal semantic misalignment, limiting sample efficiency and decision consistency. We present NaviVLN, a dual-system policy built on THREAD (Task-centric Hierarchical Representations via Embodied Alignment for Dual-system), aligning vision, language, and trajectories through data, model, and training. THREAD fuses navigation information into mutually informative visual and textual targets grounding instruction semantics and task progress. Joint Attention integrates implicit visual and textual chain-of-thought during training, shaping shared representations through auxiliary and navigation supervision. Staged training connects these representations to a local trajectory expert; deployment discards auxiliary decoders, using asynchronous intention querying for efficient execution. On R2R-CE Val-Unseen, NaviVLN achieves approximately 10% higher SR than Uni-NaVid using roughly 10% of its reported training data. Ablations show absolute SR gains of 8.49% and 4.30% over action-only and unaligned multimodal baselines, respectively, supporting the effectiveness of full-pipeline task alignment. Further gains with increasing training data suggest strong scaling potential.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.