acceptodds
Under review as a conference paper at ICLR 2027

How Sound Meets Language: Causal Mechanisms of Audio-Language Integration in Large Audio Language Models

Abstract

How do large audio language models (LALMs) integrate task-relevant auditory evidence into language computation? We introduce AudioEvidence, a controlled testbed with temporal evidence annotations, matched non-evidence controls, and counterfactual sound replacements, to trace this process across nine LALMs. At the backbone input, the causal effects of auditory evidence remain temporally localized. Within the backbone, eight models share a depth-wise organization: auditory evidence states exert stronger effects early, direct evidence access is most consequential at intermediate layers, and text-state effects strengthen with depth. Joint interventions establish that blocking intermediate evidence access weakens the influence of subsequent text states on the answer. The critical access stage also supports additional tasks and longer audio inputs. Head contributions vary in concentration across models, with MiMo-Audio exhibiting sparse organization that parallels retrieval-head findings in text LLMs. High evidence-attention scores alone do not establish causal importance. Finally, amplifying selected attention heads in MiMo-Audio improves audio hallucination detection and reduces task deviations in speech recognition. These findings provide mechanistic priors for more efficient audio scaling and modality integration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.