acceptodds
Under review as a conference paper at ICLR 2027

Quantizing Rhythm with Music Foundation Models

Abstract

Rhythm quantization, mapping the expressive timings of a musical performance into a score's precise note values, must be learned from fewer than a thousand paired performances and fewer than two hundred scores. We show that a pretrained symbolic-music foundation model can be adapted into a rhythm quantizer with no new parameters: fed the performance as its anticipated control stream, the Anticipatory Music Transformer becomes a fixed-lookahead transcriber of entire songs, fine-tuned on interleaved pairs of events from a performance and its corresponding score. Musical pretraining substantially improves this adaptation. Under an identical training recipe and checkpoint-selection rule, the pretrained model more than doubles whole-song strict note F1 relative to random initialization (8.7% vs. 3.2%). We further find that teacher-forced fine-tuning in this data-limited regime is susceptible to compounding errors during whole-song autoregressive inference. To address this, we introduce a three-stage training curriculum that progressively bridges the distribution shift between teacher-forced training and model-generated inference. On the 59-performance ASAP test set, our model reaches 9.33% strict note F1 and 77.93% inter-onset-interval F1, compared with 8.96% and 67.81%, respectively, for the strongest task-specific baseline. These gains are accompanied by a substantially higher rate of producing outputs at the reference score's time scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.