SLS: Structured Local Supervision for Fine-Grained Speech-Text Alignment
Abstract
Fine-grained speech-text alignment, such as understanding subtle shifts in speaking rate or vocal delivery, requires audio-text models to assign high similarity to accurate descriptions, while distinguishing them from inaccurate ones. We introduce Structured Local Supervision (SLS), a framework that adapts pretrained audio-text models using structured captions with SigLIP and supervised contrastive learning. Accurate captions encourage invariance to paraphrasing and partial descriptions, while inaccurate captions define targeted attribute contrasts and guide in-batch audio-negative mining. Across three backbone families, SLS substantially improves speech retrieval, classification, and controlled fine-grained ranking. Notably, SLS fine-grained alignment extends beyond retrieval, offering a lightweight, zero-shot proxy for instruction-conditioned text-to-speech evaluation. SLS rivals multi-billion-parameter audio LLMs in correlation with human adherence ratings. The findings highlight SLS as both a general recipe for attribute-sensitive representations and an efficient, scalable foundation for generative speech evaluation and alignment objectives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.