Rethinking Audio Visual Segmentation with Asymmetric Visual Entity Construction and Audio Querying
Abstract
Audio-Visual Segmentation (AVS) requires associating unlocalised audio cues with spatially dense visual entities. Prevailing methods largely rely on symmetric modality fusion, which can entangle modality-specific representations and induce spurious audio-visual correlations. Such dense coupling overlooks a crucial inductive bias of AVS: vision determines what entities are present and where they are located, whereas audio primarily provides evidence about which visible entities are acoustically active. Motivated by this asymmetry, we reformulate AVS as an asymmetric entity-querying problem and propose AEQNet that separates visual entity construction from candidate-grounded acoustic verification through four complementary components. Firstly, we propose the Hierarchical Visual Candidate Constructor (HVCC), which uses adaptive aggregation of different visual features and mask-conditioned candidate abstraction to construct compact object-centric visual candidates, addressing the spatial ambiguity of direct dense audio-visual interaction. Secondly, we propose the Coarse-to-Fine Audio Query Router (CF-AQR), which uses coarse semantic matching and candidate-constrained fine verification to progressively determine the audio activation of visual candidates, addressing visual ambiguity and simultaneous multi-source sounding. Thirdly, we propose the Factorised Semantic Composer (FSC), which explicitly combines spatial membership, semantic identity, and audio activation to generate structured semantic predictions, avoiding entangled multimodal decoding. Finally, we propose Progressive Candidate-Grounded Counterfactual Query Learning (PCQL), which uses progressive audio-query learning and hard audio counterfactuals to enforce genuine audio-dependent candidate decisions, reducing visual shortcuts and spurious audio-visual correlations. Extensive experiments on standard AVS benchmarks demonstrate the effectiveness and generality of our framework and show superior performance over existing methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.