acceptodds
Under review as a conference paper at ICLR 2027

Specialist Foundation Models as Tools for Generalist Navigation

Abstract

Generalist navigation requires diverse perceptual and reasoning capabilities to support tasks such as instruction following and object search. Recent generalist navigation models typically learn these capabilities within an end-to-end policy, so adding or improving individual capabilities can require further training of the navigation model. We argue that a navigation decision model need not master every capability itself; it can instead use independently pretrained specialist models as external tools to expand navigation capabilities without retraining the decision model. We present Perceive–Imagine–Verify (PIV), a loosely coupled framework that augments a shared decision model trained across navigation tasks with specialist foundation models as external tools. Perception tools recover missing depth and reconstruct panoramic context, foresight tools generate future observations conditioned on candidate moves, and verification tools ground targets to validate the decision model's stopping proposals. This separation offers two benefits. First, compatible tool backends can be replaced or upgraded while keeping the decision model fixed, providing a path to benefiting from advances in specialist models. Second, perception tools enable one checkpoint to operate across panoramic and forward-facing RGB and RGB-D configurations without sensor-specific fine-tuning. PIV achieves state-of-the-art performance in both vision-and-language navigation and object-goal navigation with one shared decision model. Specifically, it attains success rates of 71.6% on R2R-CE val-unseen and 56.6% on HM3D-OVON val-unseen.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.