Adaptive Relation Weighting and View Curriculum for Supervised Multi-Label Cross-Modal Hashing
Abstract
Cross-modal hashing maps heterogeneous data into compact binary codes for efficient retrieval. Existing methods mainly assume fixed image/text input configurations, while practical systems may need to jointly handle paired, image-only, and text-only inputs. This unified setting introduces heterogeneous and dynamically changing optimization difficulty across the nine directed relations among these input conditions. To address this issue, we propose a unified supervised multi-label cross-modal hashing framework that couples adaptive directed relation weighting with an adaptive view curriculum. Relation weights are estimated from raw contrastive statistics to calibrate multi-label contrastive learning, while the learned relation states further guide the sampling of paired, image-only, and text-only views for subsequent training. Experiments on COCO, NUS-WIDE, and MIRFlickr25K demonstrate the effectiveness of the proposed framework across multiple hash-code lengths and eight directed retrieval protocols.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.