Do Omni Models Use How We Speak? Diagnosing and Improving Paralinguistic Understanding
Abstract
Strong performance on speech and multimodal benchmarks does not establish whether omni models use how an utterance is spoken to guide their decisions. We investigate this question through controlled translation selection, asking models to choose the translation better synchronized with the audio while varying vocal expression, holding the transcript and candidate translations fixed, and testing both candidate orders. Evaluated models, including commercial closed-source models, often fail to adjust their choices correctly as vocal expression changes, and most choices persist even when speech is replaced with silence, indicating limited use of vocal information beyond the transcript. We introduce Paralinguistic Counterfactual Training (PCT), which combines same-transcript audio pairs with opposite target choices and human translation-quality preferences, using candidate-order augmentation and supervised low-rank adaptation. On Qwen2.5-Omni-3B and 7B, PCT increases paired correctness from 2.65% to 49.29% and from 8.17% to 50.71%, respectively, with gains also observed in four further post-trained omni models. Additional evaluations show transfer to general audio understanding and paralinguistic tasks, and largely preserved performance on other evaluated input modalities for Qwen2.5-Omni. These findings show that explicitly linking vocal expression to correct decisions can improve audio use beyond the supervision task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.