MuseAdapt: Melody-Preserving Music Adaptation with Representation and Reward Alignment
Abstract
Melody-conditioned music generation provides fine-grained control over musical content, but adapting a reference recording to new musical conditions while faithfully preserving its melody remains challenging. Existing approaches often derive melody conditions from the same audio used as the generation target, causing the conditioning representation to retain source-specific acoustic characteristics and reducing its robustness to cross-style adaptation. To address this problem, we introduce MAMA-20k, a large-scale real-world dataset containing melody-aligned music pairs with shared melodic content and diverse acoustic realizations, and propose MuseAdapt, a framework for melody-preserving music adaptation through representation and reward alignment. For representation alignment, we train a convolutional–Transformer encoder to learn temporally structured melody representations from aligned music pairs and use them to control a pretrained flow-based music generator via ControlNet. For reward alignment, we construct preference pairs from complementary melody, key, and text-alignment rewards and further optimize the generator with direct preference optimization (DPO). Experiments show that MuseAdapt achieves substantially stronger target-condition adherence than reconstruction-oriented baselines while maintaining comparable melodic consistency, yielding a better balance between melody preservation and musical adaptation without sacrificing generation quality. Audio examples can be found at https://museadapt.github.io/}{https://museadapt.github.io/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.