Conditional Causal Learning for Multimodal Sound Source Localization
Abstract
Multimodal sound source localization identifies image regions responsible for an input sound without spatial annotations. Most self-supervised methods treat synchronized audio-visual observations as positives and mismatched pairs as negatives. However, synchronized modalities encode both the sounding event and recurrent scene context, allowing pair discrimination to exploit contextual regularities without identifying the actual sound source. This issue is further complicated by modal asymmetry: audio captures event content without image-plane position, whereas vision preserves spatial structure together with latent appearance and acquisition biases. We propose Conditional Causal Learning (CCL), which formulates causal adjustment as a conditional design spanning representation learning and contrastive optimization. Rather than applying a uniform correction to heterogeneous biases, CCL selects conditioning information according to bias observability and its role in the learning process. Specifically, CCL suppresses context-dependent acoustic variation, preserves audio-consistent spatial information across alternative visual contexts, and reduces contextual bias in contrastive pair construction. Coordinating these adjustments limits shortcut dependence while retaining the semantic and spatial information required for localization. Extensive experiments across multiple benchmarks and perturbation settings demonstrate the effectiveness and robustness of CCL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.