acceptodds
Under review as a conference paper at ICLR 2027

Think-to-Act: Grounded Reasoning Supervision for Generalizable Vision-and-Language Navigation

Abstract

The dominant recipe for continuous vision-and-language navigation (VLN) fine-tunes a pretrained vision-language model (VLM) with an action cross-entropy loss. This loss rewards the correct action but places no semantic constraint on the model. The policy therefore overfits to action and pixel-goal outputs and loses the pretrained semantics that generalisation relies on. We call this failure action overfitting. It goes unnoticed in training environments but hurts navigation in unseen ones. We present Think-to-Act (T2A), which uses structured reasoning as a training-time semantic anchor. At every decision point, T2A supervises four types of navigation semantics (scene understanding, instruction progress, decision rationale, and stop condition). The supervision comes from T2A-GR, a grounded reasoning dataset that covers nearly the entire R2R-CE training split. An action-first target layout prevents teacher-forced leakage. Reasoning dropout and a path-consistency objective train the reasoning-free path used at deployment, so no reasoning token is decoded at test time. Under matched training, T2A improves val-unseen success rate by 5.7 points on R2R-CE and by 4.8 points on RxR-CE, whose instructions are never annotated. It is also more effective than data replay and single-field auxiliary objectives under the same protocol. Trained only on three simulated datasets, the reasoning-free semantic system transfers zero-shot to a quadruped robot in unseen physical scenes. It succeeds in 86.7% of trials, against 38.3% for the InternVLA-N1. Code and data will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.