FAME: Multi-Level Focus-Aware Mixture-of-Experts with Auxiliary Correctness Supervision for Audio Reasoning
Abstract
Audio reasoning has been a challenging task for Large audio-language models (LALMs) due to modality imbalance where models disproportionately rely on textual cues and under-utilise audio information. This stems from typical LALMs' architectures which process audio through a single and entangling heterogeneous aspect of audio understanding into one shared representation via an audio encoder and then align it with a text language model to perform downstream tasks. This limits both audio feature encoding and the model's ability to selectively attend to the aspect of audio representation most relevant to the task. We introduce FAME, a hierarchical focus-level framework trained through a progressive three-stage curriculum. At the low level, five aspect expert heads followed by on audio encoder specialised on different aspect-labelled corpora, independently capture distinct facets of audio evidence, and a question-driven soft Mixture-of-Experts router is trained to route and blend their features. At the mid level, focus-augmentation names the relevant experts within the input question to sharpen this routing, reinforced by an auxiliary correctness objective that teaches the model to reconsider uncertain answers. At the high level, focus-augmentation further expands the model's attention to abstract, notion-level concepts beyond individual experts. Crucially, none of FAME's training data overlaps with the audio reasoning benchmarks used for evaluation, making our evaluation fully out-of-domain. Across MMAU-Pro, MuCho Music, Clotho-AQA, and ADQA, FAME improves accuracy by 2.81% over the frozen backbone, which relies on implicit, emergent expert specialisation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.