UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-resource Music Understanding
Abstract
Recent advances in large audio-language models (LALMs) have unlocked impressive capabilities in music understanding and reasoning. However, current LALMs exhibit severe performance degradation when applied to diverse global musical traditions, particularly low-resource folk music rooted in rich cultural contexts. Due to significant data imbalance and the absence of dedicated evaluation protocols, existing models often fail to capture the nuanced structural, acoustic, and stylistic characteristics of local genres. To bridge this gap, we present UniVerse, a comprehensive, reproducible framework for culturally-aware low-resource music understanding. UniVerse consists of two core components: (1) UniVerseBench, an expert-guided benchmark featuring 5,042 meticulously curated Q&A pairs across 38 distinct cultural entities to evaluate fine-grained musical comprehension; and (2) UniVerseSet, a fully automated multi-turn dialogue dataset synthesized to facilitate multimodal alignment. Leveraging UniVerse, we present the first systematic study investigating multimodal imbalance learning strategies across both Dense and Mixture-of-Experts (MoE) LALM architectures. Extensive experiments demonstrate that our automated curation paired with latent representation alignment yields significant gains on dense backbones (+14.9 pp), while standard SFT remains a resilient baseline on sparse MoE architectures. Nevertheless, diagnostic evaluations reveal a remaining gap between surface-level textual alignment and deep acoustic comprehension, highlighting key challenges of building general LALMs for culturally inclusive music. Our anonymous demo page is available here.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.