KG-ReflectGround: Evidence-Guided Zero-Shot 3D Visual Grounding
Abstract
Zero-shot 3D visual grounding (3DVG) requires reliable spatial reasoning without task-specific training, yet achieving both strong grounding performance and computational efficiency remains challenging. Existing LLM-based approaches generally incur lower inference costs but remain limited in capturing visually critical cues, whereas MLLM-based approaches provide stronger visual disambiguation at the expense of substantially higher computational overhead. To reconcile these competing requirements, we propose KG-ReflectGround, an efficient zero-shot 3DVG framework that adaptively allocates spatial reasoning and visual verification according to evidence gain. The framework constructs a query-aware sparse spatial knowledge graph over candidate 3D object instances, prioritizing discriminative relations to provide compact, query-relevant structural evidence. Relation-ordered evidence-gain reasoning progressively refines the candidate set and adaptively halts when additional relational evidence provides limited benefit, thereby reducing reasoning cost. An evidence-aware verification policy further assesses whether the accumulated structural evidence is sufficient and selectively invokes MLLM-based visual verification only for ambiguous or appearance-dependent queries. Experiments on ScanRefer and Nr3D demonstrate a favorable balance between zero-shot grounding performance and multimodal computation. KG-ReflectGround achieves 52.0% Overall [email protected] and 47.8% Overall [email protected] on ScanRefer, with an average inference latency of 3.48 s per query.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.