acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Multi-Word Anchoring Network for Fine-Grained Image-Text Matching

Abstract

Image-text matching bridges vision and language, which is a fundamental task in multimodal intelligence. The key challenge lies in establishing accurate alignment between two heterogeneous modalities. Existing methods typically rely on either dense soft attention, which suffers from cross-modal noise, or rigid hard assignment, which causes semantic information loss by artificially restricting region-word correspondences to one-to-one mappings. To address the inherent semantic density imbalance between visual and textual modalities, we propose an Adaptive Multi-Word Anchoring Network (MWAN). Specifically, we introduce an Anchored Alignment Mechanism (AAM) as a theoretically motivated adaptive gating approach tailored for cross-modal matching. By utilizing context-aware dynamic thresholds, AAM effectively grounds each image region to multiple semantically relevant words, achieving robust many-to-many alignment. Moreover, to overcome the distinguishability bottleneck in top-ranked retrieval results, we design a Residual Semantic Re-ranking (RSR) module that explicitly isolates and exploits previously suppressed residual information to discriminate among semantically similar candidates. Extensive experiments on Flickr30K and MSCOCO demonstrate that MWAN achieves superior performance. Notably, RSR also serves as a universal plug-and-play refinement module to enhance existing methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.