DID: 2D Autoregressive Modeling with a Decoder-in-Decoder Architecture for Symbolic Music Generation
Abstract
Symbolic music generation aims to create compositions from discrete musical events. Existing autoregressive methods flatten multi-bar events into a single sequence, losing the explicit 2D structure, while non-autoregressive approaches often rely on fixed grids, limiting flexibility for variable-length generation. In this work, we propose DID, a 2D autoregressive framework that models musical events across the time and note dimensions.This design preserves the two-dimensional organization of music while retaining autoregressive flexibility. DID follows a Decoder-in-Decoder architecture with two decoders: the time decoder updates a running musical history along the time dimension, while the note decoder uses this memory to generate events along the note dimension. To improve harmonic consistency, we introduce harmonic Plan events as intermediate tokens in the note decoder's output sequence, turning the given conditions into a soft harmonic constraint on the notes that follow. Experiments on multiple public datasets show that DID improves harmonic and rhythmic coherence over prior methods, while supporting flexible generation scenarios, including variable-length music continuation, chord-conditioned piano generation, and accompaniment generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.