ErrandVLN: A Benchmark for Intent-Driven Long-Horizon Continuous Vision-and-Language Navigation
Abstract
Existing continuous vision-and-language navigation (VLN) benchmarks primarily evaluate single-goal route following, whereas everyday household requests may require agents to visit multiple semantic landmarks in a prescribed order. This mismatch leaves **intent-driven, multi-stage navigation** insufficiently characterized. We introduce **ErrandVLN**, a benchmark that pairs continuous trajectories with diverse role-conditioned requests, allowing the same ordered landmark requirements to be expressed across varied activity contexts. Constructed on InteriorGS environments in NVIDIA Isaac Sim through an automated semantic-geometric pipeline, ErrandVLN comprises 913 scenes, 65,134 unique physical trajectories, and 502,233 instructions, with evaluation splits for unseen scenes, novel trajectory compositions, and unseen instruction styles. To address compounding errors over ordered stages, we propose **VLN-JEPA**: a continuous trajectory proposer (TP-Flow) generates candidate motions, an action-conditioned latent world model (NaviLat) predicts future latent states under these candidate motions without decoding future observations, and a compact 2.35B VLM controller uses these predictions together with the instruction and observation history to generate the next executable action. On ErrandVLN Test Scene-Unseen, VLN-JEPA achieves superior performance across multi-stage milestone completion (20.5% GCR, 16.1% GC-SPL, 29.2% OSR) and competitive task success (12.6% SR), while substantially reducing physical collisions by 40% compared with DualVLN. On a single NVIDIA RTX 4090, it achieves 94 ms per policy-inference step (10.7 FPS) and 5.3 FPS in complete closed-loop execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.