acceptodds
Under review as a conference paper at ICLR 2027

MuseDuo: Unified Audio-MIDI Music Generation, Separation, Transcription, Rendering, and Editing

Abstract

Music generation has developed as two largely separate fields: symbolic generation, which produces MIDI or scores, and audio generation, which models waveforms. Musicians work across both, composing and arranging in symbolic terms and working directly with timbre and texture in the audio domain. We present MuseDuo, a diffusion/flow transformer (DiT) that is 1) trained on both music mixes and multi-track instrument stems, and 2) jointly models audio and MIDI, enabling a broad range music tasks. MuseDuo comprises three key mechanisms: (i) a unified audio-MIDI multimodal latent rectified-flow architecture with a novel MIDI-VAE, (ii) a concise training recipe that realizes diverse tasks via systematic inference-time configuration of input, output, and conditioning, and (iii) on-the-fly MIDI extraction that alleviates the need for costly annotation. We fine-tune a pretrained text-to-audio music DiT, and evaluate on 10 music tasks spanning text-to-audio, text-to-MIDI, multi-stem audio generation, open-vocabulary source separation, transcription, MIDI-to-audio rendering, timbre transfer, and note-level editing. With a unified design, MuseDuo matches or outperforms 12 recent task-specific baselines on the majority of evaluated tasks and metrics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.