acceptodds
Under review as a conference paper at ICLR 2027

OmniMapNav: Map-Guided Multimodal Goal Navigation via Dual-Grounding Map Reading and Lightweight Semantic World Modeling

Abstract

Visual navigation should support intuitive goal specification and efficient adaptation to unfamiliar environments. However, PointNav, image-goal navigation, and conventional route-following vision-language navigation require coordinates, reference images, or route descriptions that users may not readily provide; ObjectNav category labels leave ambiguity among target instances. Without a global spatial prior, inferring unfamiliar layouts from partial observations can require extensive exploration in large, multi-room environments. To bridge this gap, we introduce OmniMapNavBench, where one overhead RGB map serves as an interactive goal interface and a navigation prior. Its 50,241 episodes across 4,724 scenes pair map-selected points, hand-drawn route sketches, and language grounded in map–room, map–object, and object–object relations for evaluating individual cues and their combinations. Effective map utilization requires grounding local observations and heterogeneous goals in this common spatial frame. To this end, we propose OmniMapNav: its Dual-Grounding Map Reader estimates an agent pose-heading belief and a goal anchor field to retrieve agent- and goal-relevant map tokens. Furthermore, a Semantic Foresight Predictor estimates one-step, action-conditioned next-observation features; a semantic foresight scorer combines these with grounded context to refine action scores. OmniMapNav achieves an average geodesic success rate of 37.31% at 3 m versus 26.22% for the strongest baseline, with gains across all four goal conditions and stricter thresholds. Real-robot experiments demonstrate map-based goal specification and closed-loop navigation. Code is available at https://anonymous.4open.science/r/omnimapnav-9DDF.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.