Learning Not to Discard Prematurely: Trajectory Learning for Interactive 3D Visual Grounding
Abstract
3D visual grounding aims to localize the query-matched target object in a 3D scene. In real-world scenarios such as human–robot interaction, human queries are often ambiguous and may be compatible with multiple candidate objects, necessitating iterative interaction to identify the intended target. However, as the interaction progresses, candidate filtering based solely on current query–object matching often fails to explicitly account for its consequences for subsequent interaction. We observe that interaction trajectories reveal which early filtering decisions successfully retain the target until disambiguating evidence arrives, and which make later clarification ineffective. Building on this observation, we formulate interactive 3D visual grounding as sequential candidate refinement, and propose a trajectory learning framework that uses ordered candidate transitions, subsequent feedback, and final outcomes to optimize early candidate filtering decisions. Through trajectory-level preference optimization, the framework learns to retain plausible candidates until sufficient clarification evidence arrives. To effectively propagate trajectory-level supervisory signals to the early candidate generation, we introduce a trajectory-aware adapter that integrates interaction history and candidate belief to produce residual updates for candidate representations and grounding scores. We evaluate our framework on AmbiRefer3D, an interactive 3D visual grounding dataset, and construct a more challenging subset requiring multiple rounds of clarification to uniquely identify the target. Experimental results demonstrate that our method improves candidate retention and downstream interaction performance, outperforming the direct-training baseline by an average of 20.56% in TS and 3.66% in Acc across grounding backbones. These results highlight that learning how to avoid prematurely discarding candidates is crucial for achieving reliable interactive 3D visual grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.