acceptodds
Under review as a conference paper at ICLR 2027

Rethinking, Revealing, and Reducing Real Modality Gap of Audio Language Models under Linguistic-Paralinguistic Trade-Off

Abstract

Large Audio Language Models (LALMs) perform worse on audio than on semantically equivalent text, which is known as the modality gap. Recent studies claim that the gap can be reduced without degrading the pre-trained paralinguistic capability. However, we argue that the current protocol, which measures the two sides on disjoint data respectively built for linguistic- and paralinguistic-centered tasks, ignores their real-world tension and thus cannot actually support the claim. In response, we propose ReMG, a new protocol that crosses two linguistic contents with two paralinguistic states into a matched 4-tuple, holding the remaining acoustic covariates fixed. Each audio clip then has a variant that changes along only one axis and is also queried on both axes. This leaves no shortcut for LALMs but to directly trade off linguistic against paralinguistic capability on the same audio, only under which we believe a real modality gap can be meaningfully claimed. The result under ReMG challenges the previous view: paralinguistic collapse is widespread, and the only released checkpoint shows no substantial gap reduction. ReMG helps address this situation by guiding LALM fine-tuning. We use simple on-policy distillation without stronger teachers or ground-truth labels, so as not to mix the protocol contribution with that from new knowledge learning, as existing methods do. Across nine LALMs, four baseline methods, and seven benchmarks, the proposed ReMG is demonstrated to effectively reveal and reduce the real modality gap under the linguistic-paralinguistic trade-off.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.