GEAR: An Evidence-Adaptive VLM Agent for Zero-Shot 3D Visual Grounding
Abstract
Zero-shot 3D visual grounding aims to localize objects from natural-language descriptions without task-specific 3D training. Existing approaches largely follow two paradigms: proposal-based methods enable efficient and reliable geometric reasoning but cannot recover missing objects, while RGB-D reconstruction provides stronger visual recoverability at substantially higher computational cost and with additional sources of error. To bridge this gap, we introduce GEAR, an evidence-adaptive ReAct-style agent that unifies proposal-based grounding, selective verification, and on-demand candidate repair. Starting from proposal hypotheses, GEAR adaptively selects and revisits grounding operations based on the query and evolving evidence, invoking semantic or visual verification for ambiguous hypotheses and bounded RGB-D reconstruction only when candidates are missing or unreliable. A candidate-bound memory maintains consistent evidence across grounding operations, while a role-aware geometric solver handles compound ordinal and local-anchor expressions. GEAR achieves 60.4% [email protected] and 52.4% [email protected] on ScanRefer and 56.0% accuracy on Nr3D, surpassing TAB, the previous state-of-the-art method, while reducing token usage and inference time by approximately 50%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.