Learning Navigation Corrections from a Weak Recovery Policy
Abstract
Learning to recover from navigation errors can disrupt actions that an agent already performs well. Yet a recovery policy that is unsuitable for deployment may still offer useful local supervision. We propose RCSR-D, a training procedure that separates the usefulness of a policy's action proposals from its ability to navigate complete trajectories. A frozen recovery policy supplies candidate turns. An offline rule selects proposals under empirical retention criteria, and a student initialized from the original policy learns the selected first actions while matching reference behavior on the remaining states. Parameter interpolation moderates the learned update before deployment. On the same reserved 128-episode cohort in the navigation-tuned 9B setting, the original and recovery policies achieve 33.59% and 6.25% success, respectively, while RCSR-D students achieve 42.97–51.56% across three training seeds. On the full 1,839-episode R2R-CE validation-unseen split, the primary student improves success from 37.25% to 46.93% and success weighted by path length from 29.36% to 36.28%. These results show that a recovery checkpoint can remain a useful training resource despite poor navigation performance. RCSR-D provides a way to convert its local proposals into improved navigation while deploying a single student.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.