Grounding Visual Evidence for Long-Range Language-Conditioned Navigation
Abstract
Long-range language-conditioned visual navigation (LCVN) requires embodied agents to follow natural-language instructions across environments spanning hundreds of meters. Topology-based LCVN relies on visual appearance for target retrieval, map construction, and navigation, but appearance becomes an unreliable proxy for task relevance at long range: similar landmarks confuse retrieval, visually redundant observations can be instruction-critical, and appearance variation causes waypoint errors to accumulate. We present **GRONav**, a unified grounded topology-based navigation framework that makes visual evidence decision-specific: retrieval is grounded in action and trajectory context, map compression in semantic composition, and navigation in geometric structure distilled from monocular depth. Experiments in CARLA and on a real quadrupedal robot demonstrate consistent gains in retrieval accuracy, navigation success, and robustness over strong baselines, while maintaining the lowest inference latency. GRONav achieves 91.7% target retrieval accuracy, reduces topology size by 68.0%, and improves real-world navigation performance by more than 25.0% on 60–200m tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.