The Goal Is Not Enough: Vision-and-Language Navigation under Negative Spatial Constraints
Abstract
Vision-and-Language Navigation in Continuous Environments (VLN-CE) is typically evaluated by whether an agent reaches the goal, rather than whether it obeys what the instruction forbids. Consequently, an agent may reach the target while entering a restricted region or approaching a fragile object—an unsafe success that standard VLN metrics fail to expose. We introduce , a benchmark for constraint-compliant navigation under language-specified negative spatial constraints. NS-VLN grounds these constraints as forbidden regions in continuous environments and provides constraint-compliant reference trajectories, a three-level taxonomy of separated, linguistically integrated, and semantically entangled constraints, and safety-specific metrics that distinguish goal success from safe success. It further includes forbidden–allowed counterfactual pairs for diagnosing whether an agent understands how a referent is used in the instruction or merely associates particular object categories with danger. Our evaluation reveals substantial unsafe success across representative VLN-CE agents and shows that naive training on constraint-augmented data still leaves a substantial gap in reliable constraint compliance. As a complementary method for graph-based VLN agents, we introduce (Constraint-Conditioned Navigator), which learns linguistic-role-sensitive transition risks, composes them into route-level risk, and performs risk-aware waypoint selection. Together, NS-VLN and C2Nav establish negative spatial constraints as a first-class objective for continuous VLN evaluation and learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.