SMILE: Section-Aware Music-Driven Lyric Generation
Abstract
Music-driven lyric generation seeks to create lyrics that are semantically coherent across a song and aligned with its musical phrasing and section structure. A song's sections, such as verses, choruses, and bridges, organize where lyrical ideas develop, recur, and change, making them an effective interface between musical structure and lyric generation. However, existing approaches often rely on externally supplied lyric conditions or extract musical cues without inferring sections to organize the complete lyrics. To connect these levels of organization, we introduce SMILE (Section-aware Music-drIven Lyric gEneration), a framework that generates full-song lyrics from a complete recording and aligned notes without user-provided themes, lyric templates, or drafts. First, a Music Section Analyzer identifies section boundaries and functions. Within these sections, a Lyric Line Planner assigns notes to lines and predicts their target syllable counts, providing a common plan for lyrical form and local musical conditioning. A Section Lyric Generator then follows this plan to produce all passage drafts, using the corresponding notes and shared semantic context from the recording. Finally, a Cross-Section Lyric Refiner uses the complete draft and optional musical relationships to coordinate passage development and repetition while preserving the line plans. In the evaluated section-level comparisons, SMILE outperforms state-of-the-art lyric-generation baselines in music-lyric association, syllable-count accuracy, and section-format adherence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.