acceptodds
Under review as a conference paper at ICLR 2027

Disentangling the Execution Gap from the Information Gap with Conditional Computation in Vision-and-Language Navigation

Abstract

Vision-and-language navigation turns a spoken instruction into a route through a real building, so a robot can act on directions in a space it has not been given as a map. In continuous environments the agent sees only first-person frames and moves through low-level steps and turns. Recent video-language navigators process image histories and emit actions. Yet these tasks still differ in their information requirements: some require finding or grounding a target, whereas others directly specify distances, angles, or a return to the starting point. A single aggregate success rate cannot tell whether a failed navigator lacks the ability to execute or lacks the information needed to determine the correct action. We route each episode by how completely its instruction already states a self-motion program. When language specifies that program, we introduce Proprioceptive Navigation (PNav), a conditional computation architecture that holds an origin register outside the frozen policy and passes the program to an external executor. Without training, PNav reaches a new SOTA success rate on NavSpace and leaves trajectories outside its takeover scope unchanged. Further closed-loop and perceptual-aid tests indicate that the tested readout, control and bypass-perception modules recover only part of the gap to a privileged-information reference. These results suggest that VLN failures have different causes: when language already specifies an action-relevant motion program, scaling the visual model alone may target the wrong bottleneck, while other instructions still need a usable goal condition.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.