AvatarVLN: Agentic World Modeling for Vision-and-language Navigation
Abstract
World models provide a promising way for vision-and-language navigation (VLN) agents to anticipate the consequences of candidate actions before physically executing them. However, imagined futures are not inherently actionable. A synthetic trajectory may span multiple instruction-relevant events, omit events that would occur, fabricate events that would not, and provide increasingly unreliable geometry over long rollouts. We present AvatarVLN, an agentic world modeling framework that turns imperfect imagined futures into actionable evidence for VLN. Rather than relying on a task-specific evaluator, AvatarVLN leverages the reasoning capability of the navigation agent itself. The agent decomposes an instruction into ordered events and tracks its progress through them. Specifically, we introduce Anticipated Progress Grounding (APG), which models navigation progress as a latent state. Rather than committing to a single progress estimate, APG weights possible progress states according to real history and imagined rollouts to estimate the reduction in remaining navigation cost. In addition, this formulation uses instruction order to accommodate omitted intermediate events while discounting unsupported progress suggested by fabricated content. Beyond semantic grounding, we introduce Uncertainty-aware Geometric Grounding (UGG), which estimates the reliability of imagined depth from accumulated motion along each rollout and selectively projects trustworthy geometry into candidate maps for spatial reasoning. The world model is trained on a navigation corpus derived from VLNVerse and transfers zero-shot to unseen VLN benchmarks. Experiments on three VLN benchmarks demonstrate the effectiveness of AvatarVLN, with ablations showing consistent gains on success rate over the base agents on R2R-CE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.