How Sound Meets Language: Causal Mechanisms of Audio-Language Integration in Large Audio Language Models
Abstract
How do large audio language models (LALMs) integrate task-relevant auditory evidence into language computation? We introduce AudioEvidence, a controlled testbed with temporal evidence annotations, matched non-evidence controls, and counterfactual sound replacements, to trace this process across nine LALMs. At the backbone input, the causal effects of auditory evidence remain temporally localized. Within the backbone, eight models share a depth-wise organization: auditory evidence states exert stronger effects early, direct evidence access is most consequential at intermediate layers, and text-state effects strengthen with depth. Joint interventions establish that blocking intermediate evidence access weakens the influence of subsequent text states on the answer. The critical access stage also supports additional tasks and longer audio inputs. Head contributions vary in concentration across models, with MiMo-Audio exhibiting sparse organization that parallels retrieval-head findings in text LLMs. High evidence-attention scores alone do not establish causal importance. Finally, amplifying selected attention heads in MiMo-Audio improves audio hallucination detection and reduces task deviations in speech recognition. These findings provide mechanistic priors for more efficient audio scaling and modality integration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.