acceptodds
Under review as a conference paper at ICLR 2027

VLN-MEM: Scene-Memory-Guided Vision-and-Language Navigation for Reliable Arrival

Abstract

Vision-and-language navigation (VLN) asks an agent to follow instructions to a goal from egocentric observations. Recent VLN systems have improved grounding, yet they still fail in two distinct modes: they fail to reach the goal, and when they do reach it they fail to stop there. Both failures share one cause: the goal is specified as a destination, a pose and a view, which says where the goal is but not how to reach it nor whether it has been reached. In this paper, we propose VLN-MEM, a scene-memory-guided VLN for reliable arrival, which supplies both conditions from a persistent scene memory: a frozen per-scene graph of places, built offline from reference trajectories, from which retrieval returns next-hop geometry for the reachability and a stored goal view for the recognition. VLN-MEM adopts zero-initialized residual fusion to feed the reachability into the planner's visual stream, and training-free stop verification to confirm arrival against the goal's appearance and proximity. VLN-MEM reaches and SR on R2R-CE and RxR-CE val-unseen in a pre-mapped, goal-specified setting, gaining SR and SPL over a training-budget-matched backbone; against stronger full-data references it matches or exceeds their SR while navigating goal-directly. A component ablation confirms that each mechanism contributes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.