acceptodds
Under review as a conference paper at ICLR 2027

MusicSTAR: Structure-Aware Temporal Attention Routing for Song Generation from a Single Long Caption

Abstract

Controllable song generation requires specifying not only a song's overall characteristics, but also how they evolve over time. A single long caption offers a natural way to express this evolution by organizing global characteristics and ordered section-level descriptions into one structured temporal specification. Yet such a description does not by itself ensure reliable temporal control: during autoregressive generation, the model can access conditions from different sections, leading to temporal condition misrouting. Motivated by this observation, we propose MusicSTAR, a structure-aware framework for temporal attention routing from a single long caption. MusicSTAR factorizes section-level conditions and dynamically routes attention to the relevant section while preserving access to acoustic history, with a temporal scheduler adaptively controlling transitions between successive conditions without explicit timestamps or durations. We further introduce local-to-global preference optimization, progressing from section-level controllability to whole-song musicality. Experiments show that MusicSTAR substantially improves fine-grained temporal instruction following while maintaining strong overall song quality, demonstrating that effective temporal control depends not only on what is specified, but also on how that information is accessed during generation. Our demo are available at https://anonymous.4open.science/w/ICLR_30705_demopage/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.