IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction
Abstract
Goal-specified navigation matches a known target description against scene objects. An implicit request can suggest several plausible targets, requiring goal interpretation alongside object search. Average success alone obscures consistency across expressions and execution of inferred goals. We introduce IntentionNav, a benchmark that evaluates expression consistency and navigation under shared goal inputs. It pairs four English expressions of each of 500 intents with the same scene, designated target, and starting pose across 176 indoor scenes. A complementary study uses a revised episode manifest to compare two navigation systems under shared category predictions and correct-category controls. Across three reference backends, 36.6–45.4% of tasks exhibit both successful and unsuccessful outcomes across their expressions, despite similar average success rates. Repeated executions provide a reference for within-expression variability. With identical category predictions, the reference system and MTU3D reach the designated target neighborhood in 60.8% and 34.6% of episodes, and finish there in 18.4% and 15.0%, respectively. Task-level consistency and separate measures of search coverage and completion reveal differences that average success obscures. Code and data are available at https://anonymous.4open.science/r/IntentionNav-D7D1 and https://huggingface.co/datasets/Anonymous260726/IntentionNav, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.