Revisiting Noisy-Label Cross-Modal Hashing with Dual-Source Semantic Consensus
Abstract
Most cross-modal hashing (CMH) methods encourage semantic consistency across modalities, implicitly assuming that cross-modal agreement is uniformly beneficial. Under noisy labels, however, different modalities can respond differently to the same supervision, making indiscriminate semantic alignment susceptible to propagating unstable predictions across modalities. This motivates a key question: how can cross-modal semantic interaction be conditionally regulated when modality-specific predictions exhibit heterogeneous reliability? To address this issue, we propose a dual-source semantic consensus framework, termed DUET, which regulates cross-modal semantic interaction from complementary temporal and instantaneous perspectives. Specifically, Temporal Consensus Learning (TCL) aggregates predictions across training epochs to characterize their relative temporal stability, while Modality Consensus Learning (MCL) exploits the concentration of current prediction distributions to adaptively determine each modality's contribution to the semantic consensus and the corresponding consensus constraint. By jointly considering these complementary cues, DUET avoids treating all modality predictions as equally influential and enables conditional semantic interaction during cross-modal learning. Extensive experiments on three benchmark datasets demonstrate that DUET consistently improves retrieval performance across a broad range of noise levels and exhibits effective adaptive behavior under heterogeneous modality-specific responses to noisy supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.