Can Speech Foundation Models Represent Meaning Through Prosody? Evaluating Production and Perception
Abstract
Modern speech foundation models can process and generate fluent speech, but spoken meaning is not determined by words alone. We ask whether these models can use prosody to convey and recover meaning when lexical content is held constant. Using syntactically ambiguous sentences whose interpretations can be disambiguated by prosodic phrasing, we evaluate the same communicative contrasts bidirectionally. In perception, speech-capable models infer intended interpretations from human productions; in production, models generate the same sentences for specified interpretations, and target-blind human listeners report the meanings they recover. Human listeners recover intended meanings reliably, whereas current speech models perform substantially worse and exhibit strong interpretation and response-position biases. Diagnostic analyses reveal distinct failure modes: models may respond to prosodic variation without reliably grounding their analyses in the acoustic signal, recover informative acoustic cues without mapping them consistently to linguistic structure, or verbalize appropriate prosodic knowledge without using it functionally. In production, measurable acoustic differentiation and naturalness likewise do not guarantee successful communication, while realizations more aligned with human phrasing tend to support better meaning recovery. Together, these results show that representing prosody requires more than preserving acoustic information or producing prosodic variation: acoustic realization, prosodic organization, linguistic structure, and communicative meaning must be functionally aligned.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.