acceptodds
Under review as a conference paper at ICLR 2027

SSLongNav: Benchmarking Repeated Completion of Structured Subtasks in Long-Horizon Instance Navigation

Abstract

In long-horizon navigation, success rates for individual visits do not reveal how reliably agents complete successive requests involving multiple object instances. We introduce Service-Structured Long-Horizon Instance Navigation (SSLongNav), which groups ordered instance visits into structured subtasks representing the navigation requirements of service-inspired requests, with shared service goals and per-visit intents. Whether a visit succeeds or fails, the agent proceeds to the next prescribed visit from its current pose. Built on synthetic HSSD and scanned HM3D scenes, SSLongNav comprises 450 episodes across 326 scenes, with 3,669 structured subtasks and 12,755 instance visits. Our primary metric, Strict Subtask Completion (SSC), measures the fraction of subtasks in which every prescribed visit succeeds. To establish a strong baseline on this benchmark, we further develop PEARS, which uses shared subtask context to guide search and combines selective evidence persistence, typed evidence arbitration, and risk-aware submission. On scene-disjoint evaluation splits, PEARS achieves SSC of 5.75% on HSSD and 13.94% on HM3D, exceeding the strongest benchmark-adapted competitors by 5.09 and 7.09 percentage points, respectively. Results reveal a substantial gap between partial progress and full completion, while documenting complete multi-visit requests after earlier failures in the same run.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.