SAYAKA: Semantic Attention Yields Audio–Concept Alignment for Fine-Grained Audio–Text Retrieval
Abstract
Fine-grained audio–text retrieval requires models to precisely associate events described in captions with their corresponding acoustic evidence. Existing fine-grained alignment methods follow two paths. One introduces explicit frame-level supervision during training, relying on fixed event taxonomies that do not fully adapt to retrieval scenarios with open-vocabulary captions. The other learns implicit alignment between events and multiple fine-grained prototypes, but requires computing similarity scores with these prototypes at retrieval time, incurring noticeable overhead compared to the single-vector inner product of standard global audio–text retrieval. We propose SAYAKA, a fine-grained audio–text retrieval method that injects concept-level alignment structure only during training and fully reduces to a standard dual encoder at inference. SAYAKA parses captions into semantic concepts via lightweight text parsing; Semantically Conditioned Audio uses each concept as a query to produce conditioned audio tokens from a shared audio-token sequence through semantic attention. Soft-Positive Audio–Concept Alignment treats other concepts within the same caption as weak positives rather than negatives, while a concept separation constraint encourages distinct concepts to capture non-overlapping acoustic evidence. An adaptive audio recomposition module fuses multiple conditioned representations into a single vector aligned with the global audio embedding, enabling fine-grained retrieval using only a single vector at inference. We further construct momentum queues at both the sample and concept levels to continuously strengthen the model's ability to distinguish fine-grained conceptual differences across samples. Case studies and experimental results demonstrate that SAYAKA achieves fine-grained concept–audio alignment and surpasses all current audio–text retrieval methods on AudioCaps and Clotho in terms of R@1, while maintaining the retrieval efficiency of global audio–text retrieval.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.