acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Autoregressive Music Modeling with Temporal Realignment

Abstract

Language models for symbolic music typically flatten events into a single sequence. When used for conditional generation, this creates a conflict between autoregressive sequence order and musical time: future conditioning information must appear early in the sequence, which can disrupt the underlying temporal relationships. We address this mismatch with model-level temporal realignment. A hierarchical grouping architecture first organizes tokens into timing-aware groups (notes and concurrent onsets). Rotary time embeddings then reassign temporal coordinates to the original musical times while leaving the input sequence and causal factorization unchanged. The resulting model can still be trained with the standard next-token objective. On both symbolic and audio-conditioned generation tasks, temporal realignment substantially improves timing prediction, yields higher coherence scores, and is preferred by human listeners over an otherwise identical model without realignment under random conditioning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.