acceptodds
Under review as a conference paper at ICLR 2027

NavDBench: A Diagnostic Benchmark for Evaluating General Instruction‑Following Abilities of Navigation Foundation Models

Abstract

Existing vision-language navigation (VLN) benchmarks primarily measure end- task success, providing limited insight into which capabilities constrain an agent when navigation fails. We introduce NavDBench, a diagnostic benchmark that combines automatically generated navigation behaviors with controlled in- terventions to characterize four capabilities essential for reliable VLN: Visual Grounding, Trajectory Planning, Multi-Modal Reasoning, and Long-Horizon Memory. Rather than assigning individual tasks to isolated capabilities, NavD- Bench suppresses competing requirements through controlled variations in geom- etry, target distribution, language formulation, and instruction structure to con- struct capability-dominant embodied evaluations. Across five representative VLN agents, we find that visual grounding remains unreliable and sensitive to target dis- tributions, geometric planning is a major short-horizon bottleneck, and the evalu- ated language variations have comparatively limited impact. Long-horizon failure reflects both accumulated local errors and difficulty maintaining task progress. Navigation-specific post-training substantially improves local capabilities across several diagnostic conditions, but these gains largely fail to propagate to au- tonomous exploration and sequential navigation. Our results highlight capability composition, rather than local competence alone, as a central challenge for reliable VLN.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.