acceptodds
Under review as a conference paper at ICLR 2027

Do Omni Models Use How We Speak? Diagnosing and Improving Paralinguistic Understanding

Abstract

Strong performance on speech and multimodal benchmarks does not establish whether omni models use how an utterance is spoken to guide their decisions. We investigate this question through controlled translation selection, asking models to choose the translation better synchronized with the audio while varying vocal expression, holding the transcript and candidate translations fixed, and testing both candidate orders. Evaluated models, including commercial closed-source models, often fail to adjust their choices correctly as vocal expression changes, and most choices persist even when speech is replaced with silence, indicating limited use of vocal information beyond the transcript. We introduce Paralinguistic Counterfactual Training (PCT), which combines same-transcript audio pairs with opposite target choices and human translation-quality preferences, using candidate-order augmentation and supervised low-rank adaptation. On Qwen2.5-Omni-3B and 7B, PCT increases paired correctness from 2.65% to 49.29% and from 8.17% to 50.71%, respectively, with gains also observed in four further post-trained omni models. Additional evaluations show transfer to general audio understanding and paralinguistic tasks, and largely preserved performance on other evaluated input modalities for Qwen2.5-Omni. These findings show that explicitly linking vocal expression to correct decisions can improve audio use beyond the supervision task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.