VoLN: Vision-Only Long-Horizon Navigation—Paradigm, Benchmark, and Method
Abstract
Long-horizon navigation in open environments requires sustained reasoning over changing visual evidence. We propose Vision-Only Long-Horizon Navigation (VoLN), a navigation paradigm for multi-stage reasoning that couples visual goal grounding, context-dependent cue interpretation, and closed-loop planning across a route. Goal views specify the destination, while onboard observations and interaction history support online decisions. We instantiate VoLN for aerial navigation through VoLN-UAV, a 7,210-episode benchmark across 17 simulated environments, combining continuous 3D flight, substantial viewpoint variation, and held-out scene evaluation. Active and passive semantic beacons create repeated demands for cue selection, providing an intermediate setting between instruction-guided Vision-and-Language Navigation and exploration in unfamiliar environments. We further present VoLN-MLLM, a two-stage visual-semantic planning framework that aligns self-supervised visual features with a structured semantic space and predicts short-horizon waypoints and stopping decisions from observation history, goal views, retrieved concepts, and proprioception. On Test-Unseen, it achieves success rates of 7.4%, 4.5%, and 1.8% on Easy, Normal, and Hard episodes, respectively. These results highlight remaining challenges in temporal evidence integration and reliable long-horizon execution. Code is available at https://anonymous.4open.science/r/VoLN-UAV-DDC6/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.