ARIADNE: Disambiguation-Guided Active Perception for Zero-Shot 3D Visual Grounding
Abstract
Zero-shot 3D visual grounding aims to locate an object in a 3D scene from a natural-language expression without task-specific training. Object-centric local views preserve photographic color and texture detail from captured RGB images, but may omit the anchor objects needed to judge spatial relations. Rendered global perspective views provide wider context, but small object scale, occlusion, and reconstruction artifacts can obscure the details needed for candidate disambiguation. We introduce ARIADNE, a disambiguation-guided active perception framework that adapts its observations to the ambiguity remaining among candidates without fine-tuning or test-time training. We first decompose the expression into visual claims about target attributes and relations to anchor objects and track which claims remain unresolved for each candidate. Our Relation-Centric Observation Construction (RCOC) provides complementary captured, stitched, and rendered views for comparing candidates and anchor objects in shared visual context. When several views are available for a comparison, Online Evidence Transport (OET) jointly matches the candidates' evidence needs to these views and selects the next observation. OET balances visibility and visual quality within a transport objective, with graph and history regularization coordinating view selection across candidates and successive comparisons. Together, these components form an adaptive observation and decision loop, with each new view guiding the next comparison toward the distinctions still needed to identify the target. Experiments on ScanRefer and Nr3D show gains of 3.5 percentage points in overall [email protected] and 4.2 percentage points in overall accuracy, respectively, over the current state-of-the-art zero-shot baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.