ReCoNav: Reliable Constrained Long-Horizon Vision-and-Language Navigation via a Determination Engine
Abstract
Conventional vision-and-language navigation (VLN) asks an agent to follow a described route to a destination. Practical missions can require multiple ordered visits, with hard constraints and preferences governing both target selection and connecting routes. Existing agents use memories, maps, or constraint-guided planning, but reliably propagating scoped requirements across interpretation, grounding, and complete-route planning remains challenging. We introduce ReCoNav, whose instruction contract models route requirements for joint target and route selection. Its Determination Engine uses scene and task rules to propagate route constraints across interpretation, grounding, and planning. Derived evidence determines choices, verifies judgments, and guides repair; vision-language models supply unresolved semantics. We construct a benchmark of 192 scene-grounded tasks across nine indoor scenes to evaluate joint target and route satisfaction. Existing methods show low satisfaction and frequent timeouts. On 44 constraint-resolving tasks without the added time limit, global scene access helps but remains insufficient: ReCoNav achieves 81.82% full instruction satisfaction, versus 29.55% for the best evaluated existing method given global scene information and 38.64% for direct planning with feedback sharing our scene information, backbone, and solver. Across four backbones, ReCoNav improves satisfaction with fewer model tokens than this feedback baseline; ablations support explicit contract-based planning, global determination, and verification with repair.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.