CIVG: Causal Inference for Accurate and Robust Generalized Visual Grounding
Abstract
Generalized visual grounding, encompassing generalized referring expression comprehension and generalized referring expression segmentation, extends classical visual grounding to multi-target and no-target scenarios. However, existing methods suffer from a critical limitation: they tend to learn superficial co-occurrence patterns induced by dataset biases, and they often fail to reliably reject no-target queries due to spurious correlations. Causal inference offers a principled solution by disentangling dataset-induced biases and enabling counterfactual reasoning about target existence. To this end, we propose CIVG, a novel framework that explicitly integrates causal inference into generalized visual grounding. CIVG introduces a Causality-Aware Deconfounding Module (CADM) that performs front-door adjustment to mitigate confounding biases from both visual and textual modalities simultaneously. We further design a Counterfactual Relevance Regularization (CFR) branch that leverages counterfactual contrastive learning to enhance the model's ability to reject no-target queries. In addition, we present CIVG-Union, a dual-branch prediction strategy that harnesses the complementary strengths of models at different scales. Extensive experiments on two tasks and four benchmarks demonstrate that CIVG achieves state-of-the-art performance, with substantial gains in both accuracy and robustness, particularly in no-target scenarios. Our results validate the effectiveness of integrating causal inference into generalized visual grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.