acceptodds
Under review as a conference paper at ICLR 2027

LyriCo: Controllable Chinese Lyric Generation with Large Language Models

Abstract

Controllable Chinese lyric generation requires models to satisfy both content requirements and fine-grained structural constraints, including section layout, line count, line length, word-group lengths, and rhyme. We present LyriCo, a structured-text framework that encodes structural specifications as control tokens and fine-tunes decoder-only language models on annotated lyrics. Its data pipeline combines classifier-based genre annotation, theme labels extracted by a large language model, text-to-speech (TTS)-inspired Chinese word segmentation, and section annotations derived from music structure analysis and lyric alignment. We introduce two lightweight mechanisms to strengthen structural control: increased loss weights on control tokens and a line-local attention bias toward relevant control tokens. On a benchmark of 500 briefs spanning Pop and gufeng, ablation studies across seven configurations at the 1.5B parameter scale evaluate the contributions of both mechanisms. Compared with the standard supervised fine-tuning baseline, the combined configuration improves line-count accuracy from 93.40% to 96.80%, section accuracy from 95.19% to 98.80%, and word-format accuracy from 97.96% to 98.52%. With a 14B backbone, LyriCo achieves higher line-length, rhyme, and word-format accuracy than prompted GPT-5.5 and DeepSeek-V4-Pro on the same benchmark, while tying with GPT-5.5 for the highest line-count accuracy. These results show that emphasizing explicit control signals can improve compliance with structural constraints in Chinese lyric generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.