Leveraging Contrastive Cues for Paralinguistic Understanding in Speech LLMs
Abstract
How Speech LLMs distinguish paralinguistic information remains poorly understood, with limited resources for studying foundational cues independently of linguistic content. We first introduce ParaCorpus, a large-scale parallel speech–descriptor corpus built using text-to-speech synthesis and verified against acoustic criteria for four foundational dimensions: speaking rate, energy, pitch, and intonation. Its paired utterances with the same words but different delivery enable controlled behavioral tests and layerwise cosine-similarity analysis of speech–descriptor matching. Results reveal strong linguistic but weak paralinguistic discrimination. Persisting disparity despite supervised fine-tuning on fine-grained annotations indicates a systematic flaw over paralinguistic understanding. We therefore propose ParaDisc, a representation-level contrastive framework jointly optimizing hard-negative speech–descriptor matching and next-token prediction. ParaDisc strengthens representation-level paralinguistic matching and behavioral discrimination, improving MMSU perception and reasoning accuracy by 2.32 and 0.78 percentage over Qwen2.5-Omni-7B, respectively. Notably, it also improves emotion recognition without emotion supervision, emerging a 4.34% gain on IEMOCAP. These findings suggest that learning to distinguish foundational paralinguistic cues through speech–descriptor matching can support paralinguistic understanding, providing an empirical basis for disentangling how these cues contribute to higher-level affective.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.