Scaling Anticipatory Music Transformers
Abstract
We seek to build capable generative models for symbolic music at the largest scale supported by the limited data available. To this end, we apply recently proposed data-constrained scaling laws to symbolic music and conduct an empirical scaling study on an improved variant of the anticipatory music transformer. The study informs the training configuration for a 1.14 billion parameter model, which achieves state-of-the-art likelihood metrics on the Lakh MIDI dataset and exhibits strong performance in human evaluation, even against closed source commercial systems like Suno v6. The model demonstrates strong extrapolative performance on a new test set of music composed in 2026, and exhibits intriguing long-context generative capabilities. However, the model's final test loss diverges from the scaling prediction, raising methodological questions about data-constrained scaling laws, at least in the symbolic music domain. We release the model checkpoint, source code, scaling law runs, and the aforementioned 2026 test set, and emphasize that all training data are fully open.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.