acceptodds
Under review as a conference paper at ICLR 2027

MUUNE: Towards Fine-Grained Full-Song Music Retrieval with Multimodal Large Language Models

Abstract

Music retrieval requires correspondence between complete songs and diverse linguistic expressions, together with recognition of multiple levels of musical attributes, including style, instrumentation, vocals, and music theory. Existing methods make substantial progress in short-clip alignment, but preserving full-song detail, understanding complex queries, and constructing high-quality supervision remain key challenges. We present MUUNE (MUsic UNderstanding Embedding), which combines music–text data construction, music knowledge adaptation, and unified representation learning with a multimodal large language model (MLLM). Audio language models and music information retrieval tools provide structured metadata for multi-view retrieval supervision and atomic counterfactual evaluation. MUUNE adapts a pretrained audio language model into a Music Reasoner and learns shared music–text representations supporting inputs of up to 300 seconds. MUUNE-EH with Music-Meta-of-Thought (MMOT) exceeds all compared baselines on all seven primary retrieval tasks, reaching 44.48% mean R@10 and 93.93% mean counterfactual accuracy, exceeding CLaMP3 by 23.30 and 15.37 percentage points, respectively. Relative to MuQ-MuLan adapted with the same contrastive data, the gains are 16.09 and 7.70 percentage points. Controlled ablations show that music knowledge adaptation and an independent embedding head improve retrieval. With the model and queries fixed, extending audio from 30 to 300 seconds improves the retrieval mean by 10.95 points. The benefit of explicit music analysis depends on the representation design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.