acceptodds
Under review as a conference paper at ICLR 2027

GhostAnchor Navigation: Execution-Updated Topological Control with Locate–Verify–Ground Target Approach for Zero-Shot Object Navigation

Abstract

Zero-shot object navigation requires an agent to find and approach a text-specified object category in an unknown environment. Vision-language models (VLMs) can infer search directions and local actions from current observations, but a one-step decision does not retain the execution history of previously selected candidates. We present GhostAnchor Navigation (GA-Nav), a training-free framework that applies VLMs to zero-shot ObjectNav through execution-updated topological control. Its Node–Ghost Topology Memory (NGTM) associates depth-derived candidate waypoints across observations and stores candidates that are not matched to Nodes as Ghosts with explicit states. After each action, NGTM updates the selected Ghost from the measured pose change and uses the resulting topology to filter the directions available to the next VLM decision. When the target becomes visible, the Locate–Verify–Ground (LVG) procedure selects and verifies a target observation waypoint, then estimates a 3D target anchor from RGB-D for approach and stopping. We evaluate GA-Nav on the complete validation sets of HM3D v0.1, HM3D v0.2, and MP3D. Among the training-free zero-shot methods compared, GA-Nav ranks first in both SR and SPL on all three benchmarks. Specifically, GA-Nav achieves absolute SR/SPL gains of percentage points on HM3D v0.1, points on HM3D v0.2, and points on MP3D, relative to the best prior training-free zero-shot result for each metric.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.