AnchorNav: Long-Horizon 3D Game Navigation with Persistent Memory and Waypoint-Conditioned Control
Abstract
Recent advances in multimodal agents have steadily improved autonomous game-playing, yet reliable long-range navigation in photorealistic 3D games remains challenging due to open terrain, extended horizons, and the absence of privileged localization. Here we present AnchorNav, a *Vision-Language-Navigation* framework built around two complementary anchors: a local visual anchor for grounding route-level decisions into short-horizon actions, and a persistent global anchor for maintaining spatial consistency over time. The local anchor projects the next navigation target into the current view as a 2D waypoint, which guides a waypoint-conditioned visuomotor controller together with the language instruction. The global anchor is provided by a persistent topological memory, where visual place recognition periodically re-anchors visual odometry to previously established locations. In photorealistic 3D games, AnchorNav effectively completes long-horizon navigation and remains competitive on RxR under the VLN-CE protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.