acceptodds
Under review as a conference paper at ICLR 2027

IP-GRPO: World-Model-Guided Reinforcement Post-Training for Vision-and-Language Navigation

Abstract

Reinforcement post-training complements supervised fine-tuning with execution feedback that helps Vision-and-Language Navigation (VLN) policies correct errors after deviating from demonstrations. However, existing RL-based VLN methods often use trajectory-level rewards that provide limited guidance on whether intermediate decisions advance instruction completion. We introduce IP-GRPO, a world-model-guided reinforcement post-training framework that converts imagined futures into progress feedback at the action-chunk level. At each decision state, an action-conditioned world model predicts the short-term visual consequences of independently sampled auxiliary action chunks. An instruction-conditioned progress scorer evaluates current and imagined states using their respective textual and visual contexts. Averaging candidate-future scores yields a prospective state-progress estimate; subtracting the current-state score converts this estimate into a dense auxiliary reward. These auxiliary rewards are accumulated along execution trajectories and combined with task-success and trajectory-alignment rewards for Group Relative Policy Optimization (GRPO). All auxiliary modules operate only during post-training, preserving the policy's original inference procedure without future prediction or candidate search at deployment. Across the R2R-CE and RxR-CE Val-Unseen splits, IP-GRPO improves success rate by 4.83 percentage points on average over the Uni-NaVid and StreamVLN base policies. The post-trained Uni-NaVid also achieves a mean success rate of 76.7% across three zero-shot real-robot tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.