acceptodds
Under review as a conference paper at ICLR 2027

LociNav: Pose-Indexed Visual Caching and Retrieval for Iterative Object-Goal Navigation

Abstract

Vision-language model (VLM) driven navigation agents process streaming egocentric video to find objects in unseen environments, yet they discard the visual features computed at each step, thus losing access to useful observations as the context budget fills. When the agent pursues consecutive goals in the same scene, this per-step forgetting wastes both the information and the computation that earlier exploration has produced. We propose LociNav, which maintains a pose-indexed store of computed visual features that grows as the agent explores. The store is organized by spatial landmarks and viewing direction, keeping it compact without discarding distinct viewpoints. When a new object goal is issued, a scoring model selects the most relevant stored observations and assembles a compact, goal-tailored memory block that is prepended to the model's input. Since the stored features need not be re-encoded and the memory block is arranged to maximize key-value (KV) prefix reuse across navigation steps, the agent consults its memory at reduced computational cost. We further construct an iterative object-goal navigation dataset and benchmark by organizing standard HM3D ObjectNav and HM3D-OVON episodes into tours of consecutive goals and incorporating the natively iterative GOAT-Bench. LociNav achieves state-of-the-art performance across the evaluated benchmarks, with its advantage growing significantly as tours deepen, while achieving over 3 Hz on-device inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.