acceptodds
Under review as a conference paper at ICLR 2027

SLS: Structured Local Supervision for Fine-Grained Speech-Text Alignment

Abstract

Fine-grained speech-text alignment, such as understanding subtle shifts in speaking rate or vocal delivery, requires audio-text models to assign high similarity to accurate descriptions, while distinguishing them from inaccurate ones. We introduce Structured Local Supervision (SLS), a framework that adapts pretrained audio-text models using structured captions with SigLIP and supervised contrastive learning. Accurate captions encourage invariance to paraphrasing and partial descriptions, while inaccurate captions define targeted attribute contrasts and guide in-batch audio-negative mining. Across three backbone families, SLS substantially improves speech retrieval, classification, and controlled fine-grained ranking. Notably, SLS fine-grained alignment extends beyond retrieval, offering a lightweight, zero-shot proxy for instruction-conditioned text-to-speech evaluation. SLS rivals multi-billion-parameter audio LLMs in correlation with human adherence ratings. The findings highlight SLS as both a general recipe for attribute-sensitive representations and an efficient, scalable foundation for generative speech evaluation and alignment objectives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.