Knowing Where to Look Next: Adaptive Regional Indexing and Evidence-Guided Search for Multimodal Lifelong Navigation
Abstract
Multimodal lifelong navigation requires robots to find successive targets specified by categories, language descriptions, or reference images in the same initially unknown environment while reusing past observations. Three bottlenecks limit effective search. First, object-indexed retrieval relies on detections made before future goals are known, carrying early omissions and mislabels into later attempts to retrieve target evidence and search cues. Second, semantic scoring ranks frontier candidates generated by geometry-only clustering of explored–unexplored boundaries, but cannot add cue-indicated observation points absent from that set. Third, scheduling must prevent new nearby geometric frontiers from repeatedly postponing pending inspections and identify worthwhile revisits when candidates are exhausted but the target remains unfound. We present TreeNav, a framework without task-specific model training that uses visual evidence not only to recognize targets, but also to decide where to search and where to look again. Its Scene Tree organizes observations into a content-adaptive regional hierarchy and retrieves raw images through regional context for goal-specific target verification and cue extraction. Regional budgets limit redundant checks while preserving verification opportunities across relevant regions. Grounded Semantic Frontiers (GSFs) turn identified search cues—goal-related objects and relevant regional information in historical and current images—into reachable observation candidates. These candidates extend historical reuse beyond direct target sightings and add key inspection locations missed by geometric frontiers during online exploration. Evidence-first wave scheduling balances discovery benefits against navigation costs, limits repeated displacement of pending inspections by new geometric frontiers, and promptly incorporates new GSFs to respond to new visual cues. After all executable candidates are exhausted, targeted recovery reuses the Scene Tree's regional index to select goal-relevant, insufficiently observed mapped locations for reinspection. In system-level comparisons on full GOAT-Bench val-unseen, TreeNav achieves 73.14% success rate and 50.14% success weighted by path length (SPL), exceeding the best reported result for each metric under the same protocol by 6.62 and 5.44 percentage points, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.