acceptodds
Under review as a conference paper at ICLR 2027

SARA: Stable Anchors and Rigorous Alignment for Fine-grained Audio-Text Retrieval

Abstract

Fine-grained Audio-Text Retrieval (ATR) requires precise token-level semantic alignment, yet existing prototype-based methods face two coupled limitations. First, without geometric constraints, learned prototypes collapse into redundant global representations, reducing fine-grained matching to global alignment. Second, mini-batch contrastive training limits the supply of hard negatives for distinguishing subtle semantic differences. We propose Stable Anchors and Rigorous Alignment (SARA) to address both jointly. To suppress prototype collapse and preserve inter-prototype diversity, SARA replaces dynamic pooling with learnable queries, applies orthogonal regularization to the resulting prototypes, and uses optimal transport to balance cross-modal prototype assignment so that every prototype contributes to matching. To further train these stabilized prototypes to distinguish fine-grained semantic differences, SARA mines hard negatives from a momentum-updated dynamic queue, focusing contrastive training on high-similarity negative audio-text pairs. Experiments on AudioCaps and Clotho under the unified encoder configuration show that SARA achieves the highest mean R@1 in both retrieval directions. On AudioCaps, SARA improves mean R@1 over the baseline (GPA) by 6.31 and 3.05 percentage points for audio-to-text and text-to-audio retrieval, respectively. Diagnostic analysis and case studies further show that SARA effectively mitigates prototype collapse and supports fine-grained semantic alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.