InfiniteVLN: Constant-Memory Streaming Vision-and-Language Navigation
Abstract
Current Vision-and-Language Navigation (VLN) models are constrained by growing context length and high computational cost in long-horizon scenarios, limiting their applicability to real-world robots that require continuous observation and low-latency decision making. In this paper, we propose InfiniteVLN, a novel streaming VLN framework that enables constant-memory navigation through a bounded, instruction-conditioned belief state. Specifically, we formalize VLN task as a Partially Observable Semi-Markov Decision Process (POSMDP) and decouple inference process into an observation step, which updates persistent visual memory, and a decision step, which generates actions based on the current memory state. To support efficient streaming inference, we introduce Asymmetric Context Caching, which stores the globally invariant instructions as a persistent instruction sink while preserving local spatiotemporal consistency within a constant-size KV space. Furthermore, we propose an Instruction-Gated Recurrent Memory module, which leverages the relative attention mass of visual tokens with respect to the instruction to construct a write-gating mechanism, and updates a long-term associative memory matrix via the Delta Rule. This design enables the model to selectively retain critical, instruction-relevant historical evidence. Our approach converts long-horizon VLN from full historical context modeling into instruction-conditioned evidence consolidation under a bounded state, which significantly reduces caching overhead and inference latency. Extensive experiments demonstrate that InfiniteVLN achieves superior performance, memory stability, and computational efficiency across both standard VLN benchmarks and challenging long-trajectory settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.