CounterGround3D: Probing and Improving 3D LMM Grounding with a Counterfactual Intervention Benchmark
Abstract
Reliable 3D grounding requires large multimodal models (LMMs) to base their predictions on the visual evidence present in the point cloud. However, existing 3D LMMs tend to be dominated by language priors and dataset-level correlations, while frequently neglecting visual input. We introduce CounterGround3D, a paired counterfactual benchmark that directly evaluates whether model predictions track controlled interventions in the point cloud. CounterGround3D comprises six types of interventions, namely, recolor, flip, add, remove, replace, and move, covering a wide range of factors in grounding tasks. Each edited scene is paired with its original counterpart, enabling evaluation of both sensitivity to answer-critical interventions and invariance to irrelevant ones. Leveraging CounterGround3D, we propose mDPO, a mixed multimodal direct preference optimization framework that employs counterfactual pairs as direct preference supervision and boosts 3D grounding with little catastrophic forgetting. Experiments across multiple widely adopted 3D LMMs show that CounterGround3D reveals substantial grounding failures that remain hidden under conventional benchmarks, while mDPO significantly improves 3D grounding capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.