acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Reasoning Context for Vision-Language Navigation

Abstract

Reasoning in vision-language navigation (VLN) can serve as an action-generation trace, auxiliary supervision, or policy context. Policy context is particularly suited to sequential navigation, making task-level guidance explicit at execution time and reusable across action decisions. Yet what this context should contain remains underexplored. Scene description, progress assessment, and subgoal planning address complementary navigation needs, suggesting that combining them should provide better guidance. Surprisingly, we find that more comprehensive reasoning context can lead to worse navigation. Through component-wise comparisons on matched navigation trajectories and action annotations, we find that subgoal planning alone outperforms the full reasoning context. However, the best way to complete a subgoal can depend on what follows it. We therefore propose Context-Think, which uses final-goal replanning to regenerate a language plan for the remaining route from the agent's current state to the original destination. Each updated plan explicitly represents the remaining subgoals, allowing future route requirements to inform current action selection while adapting to execution progress. Context-Think improves navigation success and path efficiency over both subgoal planning and full reasoning context on the widely used R2R-CE and RxR-CE benchmarks. Our findings challenge completeness as a design principle for policy context: include fewer reasoning components, but plan all the way to the final goal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.