Scalable Native Agentic Intelligence for Long-Horizon Vision-and-Language Navigation
Abstract
Vision-and-language navigation (VLN) requires agents to operate in unknown environments over long horizons while coordinating perception, memory, planning, and action. Existing agentic approaches typically orchestrate multiple specialized models, resulting in substantial pipeline complexity and inference overhead. We present *SNAIL-VLN*, a native agentic framework in which a shared multimodal language model jointly learns active exploration, action prediction, spatial perception and memory, and task planning and scheduling, with a runtime harness integrating these capabilities in a closed loop. To enable joint capability learning, we construct a spatially annotated dataset from 153 indoor scenes, including region-topology supervision and 4,742 DFS-guided exploratory expert trajectories, and curate five types of capability-specific question-answer pairs for joint training. Experiments on LH-VLN demonstrate that SNAIL-VLN achieves a new state of the art, improving the multi-stage task success rate from 2.44% to 11.40%. Ablation studies confirm the contribution of each atomic capability, while scaling experiments show that increased model capacity generally improves navigation performance, particularly on multi-stage tasks and subtask success. Code, model checkpoints, and data will be publicly released to facilitate further research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.