Long-Term Changing Grounding: 3D Visual Grounding in Continuously Varying Scenes
Abstract
Most existing 3D visual grounding tasks are conducted in static scenes that capture the scene layout only at a single instant. However, in real-world scenarios, object arrangements change over time. It is impractical to rescan the scene every time an object moves during 3D visual grounding. To address this issue, we propose a novel task, 3D long-term changing grounding, and construct a new dataset named TemScan. In the TemScan dataset, object layouts change dynamically over time, requiring models to balance perception costs while accurately comprehending the scene. Furthermore,we propose a detect-then-ground approach to address this task. Specifically, we first fine-tune a pre-trained 3D convolutional network on the ScanNet dataset to predict 3D object proposals, then build structured representations of dynamic scenes via an uncertainty-aware scene graph, and finally introduce the plan-act-evaluate reasoning module to predict the target object. In addition,we leverage LLM to distill compact and reusable skill primitives, reducing token consumption by 43% and inference latency by 18%.Experimental results on the TemScan dataset show that our method achieves accuracies of 34.98% and 34.51% under the [email protected] and [email protected] metrics, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.