StopWatch: Goal Grounding for Object-goal Navigation
Abstract
Pretrained VLMs bring rich world knowledge and semantic priors to object-goal navigation. Fine-tuning them into navigation policies has become a common recipe. Despite the progress, success rates remain low, and it remains unclear which capability the fine-tuned navigation policies lack. Our diagnosis identifies goal grounding as the dominant failure: the agent explores and remembers well, yet often fails to recognize that the goal is already in view or to estimate how far it remains. To address this failure, we let the agent ground the goal explicitly before each action, stating the semantics in view and the distance to the goal. Our experiments show that explicit grounding substantially improves success, reaching state-of-the-art results on six object-goal navigation benchmarks across three scene collections. Our analysis further shows that the policy nearly solves the task once grounding is correct, leaving the remaining headroom in perception rather than policy learning. Our work establishes goal grounding as the central problem of object-goal navigation and offers a simple yet effective way to address it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.