acceptodds
Under review as a conference paper at ICLR 2027

Towards Flexible and Controllable Music Generation for Universal Video

Abstract

Video-to-Music (V2M) generation task creates music that matches given visual content to improve audience experience. However, current methods have limitations in editing flexibility, conditional controllability, and generalizability. To meet varied user needs for general video use cases, we first construct Universe-MV, a music-video dataset for general scenarios, using large-scale web videos. We then train CueMuse, a general-purpose V2M framework supporting multi-condition control and editing, on Universe-MV. The framework includes V-Beater, a visual rhythm extraction module, and a conditional music generation backbone. Our model detects visual rhythm frames at 12 fps and can achieve precise alignment between video and music. At the same time, CueMuse can take text and editable video rhythm timestamps as inputs to satisfy additional professional requirements. Both subjective and objective experiments suggest that CueMuse achieves competitive performance and generally outperforms prior models in music-video alignment, music generation quality, and musicality. See our demo at https://anonymous.4open.science/r/demo-CueMuse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.