MusicSTAR: Structure-Aware Temporal Attention Routing for Song Generation from a Single Long Caption
Abstract
Controllable song generation requires specifying not only a song's overall characteristics, but also how they evolve over time. A single long caption offers a natural way to express this evolution by organizing global characteristics and ordered section-level descriptions into one structured temporal specification. Yet such a description does not by itself ensure reliable temporal control: during autoregressive generation, the model can access conditions from different sections, leading to temporal condition misrouting. Motivated by this observation, we propose MusicSTAR, a structure-aware framework for temporal attention routing from a single long caption. MusicSTAR factorizes section-level conditions and dynamically routes attention to the relevant section while preserving access to acoustic history, with a temporal scheduler adaptively controlling transitions between successive conditions without explicit timestamps or durations. We further introduce local-to-global preference optimization, progressing from section-level controllability to whole-song musicality. Experiments show that MusicSTAR substantially improves fine-grained temporal instruction following while maintaining strong overall song quality, demonstrating that effective temporal control depends not only on what is specified, but also on how that information is accessed during generation. Our demo are available at https://anonymous.4open.science/w/ICLR_30705_demopage/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.