acceptodds
Under review as a conference paper at ICLR 2027

Canvas-VLN: From Future Plans to Execution-Grounded Memories for Vision-Language Navigation

Abstract

Vision-and-Language Navigation (VLN) requires an agent to make long-horizon, closed-loop decisions in continuous environments by following natural-language instructions. Recent methods translate the outputs of vision-language models into executable local plans by predicting action chunks, image-space goal points, or pixel trajectories. However, existing trajectory-based approaches typically treat trajectories as disposable action outputs, while representing historical context using sequences of raw observations or textual action chunks. Consequently, future planning and past execution experience reside in different representation spaces. We present Canvas-VLN, a unified canvas–trajectory navigation framework in which pixel trajectories serve three roles: supervision targets, online plans, and navigation memory. The model predicts a future pixel trajectory on the current navigation canvas, and a local controller executes a short prefix of the predicted trajectory. The waypoints actually traversed by the agent are then reprojected onto historical anchor images to construct a Gated Execution-Trajectory Memory. This memory adaptively commits local trajectory segments according to visibility, waypoint budget, and turning events, while explicit turn states connect consecutive historical anchors. It thereby maintains a long-horizon, execution-grounded navigation context under a fixed visual-memory budget. Canvas-VLN thus establishes a predict–execute–write-back loop, in which a predicted future trajectory becomes historical trajectory evidence after execution and subsequently informs the next planning step. Trained solely on the official R2R-CE and RxR-CE navigation demonstrations, without additional training data or DAgger augmentation, Canvas-VLN achieves a Success Rate (SR) of 53.83% on R2R-CE Val-Unseen and 60.56% on RxR-CE Val-Unseen, establishing a new state of the art on RxR-CE among methods trained exclusively on task-specific navigation data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.