acceptodds
Under review as a conference paper at ICLR 2027

Which to Visit, When to Stop: Autonomous Multi-Object Search with Interactive 3D Memory

Abstract

Existing embodied navigation benchmarks support diverse goal modalities and long-horizon tasks, but typically prescribe goal sequences or task lengths, leaving two central decisions underexplored: which to visit and when to stop. We introduce MOST-Bench, a benchmark for autonomous search in embodied navigation, where agents must select target instances, organize visits, track progress, and determine when the search is complete. MOST-Bench comprises two complementary tasks: Many-Object Navigation (ManyON), which requires visiting a specified number of distinct matching instances, and All-Object Navigation (AllON), which requires finding all matching instances and deciding when to stop without knowing the target count. The benchmark contains 750 episodes across 36 HM3D scenes, covering five task–modality combinations with 2–10 targets per episode. We also introduce MOST-Nav, a training-free baseline that combines vision-language model decision-making with interactive 3D memory reconstructed from RGB observations. We evaluate seven training-based and training-free frameworks on task completion, reporting accuracy, and target coverage. The best-performing method achieves an overall success rate below 18%, highlighting the difficulty of autonomous search and motivating further research on coordinating exploration, instance-level memory, and search termination.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.