acceptodds
Under review as a conference paper at ICLR 2027

Zero-Shot Singing Style Transfer via a Structured Lyrics-Melody Bottleneck

Abstract

Zero-shot singing style transfer aims to reproduce a reference singer’s vocal identity and expressive delivery while preserving a song’s lyrics and melody. Existing conversion systems can retain unwanted traces of the original singer's style: the pitch trajectories and frame-level features used to preserve the song also encode how it was performed. We introduce , a score-conditioned, singer-prompted singing framework built around a structured lyrics–melody bottleneck. This symbolic score represents timed notes, their lyric assignments, and groups of notes sung on the same syllable (melismas), separating score-level constraints from acoustic realization. Joint lyrics–melody transcription recovers this structure from full-mix audio (vocals with accompaniment) and supplied lyrics, allowing lyric evidence to guide note selection. A reference-conditioned generator explicitly encodes this structure and uses reference audio to guide pronunciation timing and vocal expression. The framework supports both synthesis from supplied scores and conversion from recovered scores without singer-specific fine-tuning. Experiments show higher mean style-similarity and naturalness ratings for both tasks, alongside improved note transcription, note–lyric assignment, conversion intelligibility, and predicted singing quality. Audio examples are available at https://scube2027iclr.github.io/demo/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.