acceptodds
Under review as a conference paper at ICLR 2027

Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation

Abstract

To advance the development of embodied general navigation, Test-time Adaptation for Vision-Language Navigation (TTA-VLN) has attracted increasing attention, aiming to adapt pretrained policies online to previously unseen environments using only test-time observations and interaction history. However, distribution shifts in unseen environments can distort the pretrained policy's local action preferences and lead to off-course decisions. Existing methods seek to correct such deviations using test-time signals, such as predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience. Yet these signals do not directly establish whether an executed behavior actually contributes to instruction-guided progress toward the goal. Moreover, even when a test-time signal suggests a plausible corrective direction, the resulting policy change may still be unreliable and should not necessarily persist in subsequent decisions. The central challenge is therefore twofold: how to identify whether an interaction supports goal-directed improvement, and how to determine whether the resulting policy update is worth retaining. We observe that every executed action induces an immediate observation transition that exposes evidence of its local consequence. Based on this observation, we propose Credit-Guided Policy Improvement (CGPI), which uses action-induced observation transitions to recover signed, reference-relative decision credit without external outcome feedback. The recovered credit proposes a lightweight policy update, which is verified against prior credit-supported interactions and retained only when supported; otherwise, it is rolled back, while the pretrained navigation policy remains frozen. CGPI achieves consistent gains across the evaluated VLN benchmarks and navigation backbones, while qualitative robot trials further illustrate zero-shot sim-to-real feasibility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.