acceptodds
Under review as a conference paper at ICLR 2027

Grounding Visual Evidence for Long-Range Language-Conditioned Navigation

Abstract

Long-range language-conditioned visual navigation (LCVN) requires embodied agents to follow natural-language instructions across environments spanning hundreds of meters. Topology-based LCVN relies on visual appearance for target retrieval, map construction, and navigation, but appearance becomes an unreliable proxy for task relevance at long range: similar landmarks confuse retrieval, visually redundant observations can be instruction-critical, and appearance variation causes waypoint errors to accumulate. We present **GRONav**, a unified grounded topology-based navigation framework that makes visual evidence decision-specific: retrieval is grounded in action and trajectory context, map compression in semantic composition, and navigation in geometric structure distilled from monocular depth. Experiments in CARLA and on a real quadrupedal robot demonstrate consistent gains in retrieval accuracy, navigation success, and robustness over strong baselines, while maintaining the lowest inference latency. GRONav achieves 91.7% target retrieval accuracy, reduces topology size by 68.0%, and improves real-world navigation performance by more than 25.0% on 60–200m tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.