Ask What If Before Committing: Counterfactual Stability for Vision-Language Navigation
Abstract
Vision-language navigation in continuous environments often employs online topological planners to rank candidate goals on an evolving map built from partial observations. Yet a high policy score does not reveal whether a goal remains supported when the computation behind that score changes. We propose CTPNav, a framework that measures decision-time counterfactual stability by combining mean support with a penalty for response dispersion under controlled internal interventions. At the representation level, standardized edge responses gate the learned attention residual, while their dispersion adjusts the association radius for the next map update. At the action level, CTPNav probes goal competition, geometric bias, and instruction grounding to assess the stability of candidate goals. The resulting responses provide detached targets for policy learning and guide a bounded inference-time correction of action scores. On R2R-CE Val-Unseen, CTPNav achieves 61% SR and 51% SPL. On RxR-CE Val-Unseen, it achieves 54.53% SR, 45.98% SPL, and 63.19% nDTW. These results demonstrate the value of intervention-response stability for learning and topological planning. Our code is available at https://anonymous.4open.science/r/CTPNav-AD2F.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.