acceptodds
Under review as a conference paper at ICLR 2027

NavWorld: Scaling Long-Horizon Outdoor Language Navigation with Crowdsourced Street Imagery

Abstract

While language-conditioned navigation has made substantial progress in indoor environments and autonomous driving, long-horizon instruction following for ground robots in outdoor pedestrian environments remains challenging. We present NavWorld, an ecosystem for long-horizon, language-conditioned outdoor navigation built around a simple idea: automatically turning large-scale, geographically diverse street imagery into language-labeled trajectories. Our data engine sources candidate navigation trajectories from Mapillary, a large-scale crowdsourced platform for ground-level imagery spanning diverse geographic regions and capture conditions. It then filters sequences for geometric and perceptual quality and automatically generates grounded high-level instructions, intermediate subtasks, and stopping conditions from visual observations and trajectory geometry. We train a language-conditioned navigation policy that generates language subtasks from high-level instructions, tracks progress across subtasks, and predicts continuous motion from monocular egocentric observations. We evaluate NavWorld in Habitat-GS, a photorealistic outdoor navigation simulator, where it achieves a 43% success rate on NavWorld-labeled instructions, more than 9x that of the navigation Vision-Language-Action (VLA) baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.