YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Abstract
Symbolic models make musical composition explicit, while audio models generate complete recordings with composition largely implicit. We introduce YuE2, which unifies symbolic and audio music generation in one model at frontier song quality. Through symbolic planning, the model writes an editable melody-and-chord score, expands it into semantic music tokens, and renders full-song audio. The score specifies musical structure while leaving performance details to audio generation. Automatically constructed score and semantic targets let the model learn this generation process from recordings without pre-existing aligned scores. On WildSongBench, YuE2 scores 6.73 on SongBench (Wu et al., 2026) Global Avg with a two-candidate budget, exceeding all evaluated public baselines. YuE2 (best-of-8) achieves 6.96, the highest observed mean among all evaluated systems. In expert listening, YuE2 (best-of-8) receives 57.3% of overall preferences versus 30.5% for Suno v4.5 (Team Suno, 2025), and 40.4% versus 39.9% for Suno v5 (Suno, 2025). Within the same checkpoint, symbolic planning improves overall quality and musicality: 49.3% of overall judgments favor planning versus 34.6% without it. The generated audio follows the particular planned melody and harmony; score edits change the requested content while largely preserving unedited musical material. Without cover-specific training, the same checkpoint generates zero-shot covers from full scores of 948 unseen works, surpassing both evaluated cover systems on all eight work-identity retrieval metrics. This score-to-audio mapping also enables agentic music editing. An external language-model agent turns user feedback into score, style, and lyric revisions, which YuE2 renders as complete songs. Audio examples: https://yue2-0001.github.io.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.