Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation
Abstract
To advance the development of embodied general navigation, Test-time Adaptation for Vision-Language Navigation (TTA-VLN) has attracted increasing attention, aiming to adapt pretrained policies online to previously unseen environments using only test-time observations and interaction history. However, distribution shifts in unseen environments can distort the pretrained policy's local action preferences and lead to off-course decisions. Existing methods seek to correct such deviations using test-time signals, such as predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience. Yet these signals do not directly establish whether an executed behavior actually contributes to instruction-guided progress toward the goal. Moreover, even when a test-time signal suggests a plausible corrective direction, the resulting policy change may still be unreliable and should not necessarily persist in subsequent decisions. The central challenge is therefore twofold: how to identify whether an interaction supports goal-directed improvement, and how to determine whether the resulting policy update is worth retaining. We observe that every executed action induces an immediate observation transition that exposes evidence of its local consequence. Based on this observation, we propose Credit-Guided Policy Improvement (CGPI), which uses action-induced observation transitions to recover signed, reference-relative decision credit without external outcome feedback. The recovered credit proposes a lightweight policy update, which is verified against prior credit-supported interactions and retained only when supported; otherwise, it is rolled back, while the pretrained navigation policy remains frozen. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate zero-shot sim-to-real feasibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.