acceptodds
Under review as a conference paper at ICLR 2027

Mugi: Multi-task Song Generation with Joint Autoregressive-Diffusion Training and Cover-Oriented Tokenization

Abstract

Song generation has advanced rapidly in recent years, but most existing studies still focus on text-conditioned generation, with limited support for tasks such as cover song generation. Meanwhile, recent systems increasingly adopt a two-stage paradigm in which an autoregressive (AR) language model predicts high-level representations and a diffusion model renders acoustic details. However, these two modules are typically trained separately and connected only through discrete tokens. In this paper, we present Mugi, a unified song generation model that supports text-to-song, cover song, and timbre-controllable song generation within a single framework. To support high-fidelity cover song generation, we introduce a Cover-Tokenizer that represents melodic information and linguistic content as two parallel token streams, enabling the generated song to faithfully preserve melodic, harmonic, and rhythmic characteristics while maintaining the phoneme-level temporal alignment of the lyrics. Furthermore, Mugi jointly trains the AR and diffusion stages by directly conditioning the diffusion model on the continuous hidden states of the AR model, optimizing the two stages together in a single end-to-end training process using both the AR language-modeling loss and the flow-matching objective. Extensive subjective and objective evaluations demonstrate strong performance across multiple song generation tasks, with objective results competitive with commercial systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.