High-Order Interaction and Modality-Specific Distillation for Unsupervised Cross-Modal Hashing
Abstract
Unsupervised cross-modal hashing aims to encode heterogeneous data into a shared compact Hamming space without relying on semantic annotations. A key design challenge is to exploit interactions between paired image and text features during training while transferring the resulting multimodal knowledge to independently deployable unimodal hash encoders. To address this limitation, we propose a High-Order Cross-Modal Hashing framework, termed HOCH, which combines progressive cross-modal interaction with relational knowledge distillation. Its cascaded multi-head interaction blocks form products of normalized low-rank projections and use these interaction features to update the image and text representations through separate gated residual updates. Each subsequent block computes interactions from the updated representations, without explicitly constructing full outer products. An adaptive modality fusion module then combines the enhanced representations using sample-dependent weights to construct a multimodal teacher. Modality-specific relational distillation transfers both cross-modal and intra-modal similarity structures from the teacher to image and text student hash branches, enabling each branch to generate hash codes independently at inference. Experiments on MIRFlickr, NUS-WIDE, and MSCOCO with hash-code lengths from 16 to 128 bits show that HOCH achieves the best results in most evaluated settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.