Creating Music as a Go Player: a Process Reward Model of Music Generation
Abstract
Evaluating and aligning music generation systems remains challenged by the gap between static, outcome-level labels and the intrinsically dynamic, process-dependent nature of musical quality. Existing post-training methods rely primarily on coarse global rewards, restricting fine-grained policy optimization and leaving models incapable of localized error correction. Inspired by value function temporal difference (TD) learning in computer Go (e.g., AlphaGo)—where dense intermediate values are derived from sparse outcomes—we introduce MusicGo, a process reward model that assigns dense temporal rewards to partial musical sequences. By framing process reward estimation as predicting the likelihood that a prefix leads to a high-quality completion within a TD framework, MusicGo converts sparse track-level preferences into fine-grained process supervision with minimal human process annotations. Correspondingly, we release MusPR, the first dataset for training and validating process reward models in music generation. Furthermore, we evaluate MusicGo in offline and online reinforcement learning settings and present PassGRPO (Passage-level Group Relative Policy Optimization), an enhanced GRPO variant that executes policy updates at passage granularity to produce structurally coherent generations. By establishing frame-level reward diagnostics and passage-wise gradient updating, our framework lays a precedent for targeted music revision, localized inpainting, and fine-grained interactive editing. We make our audio samples available here.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.