acceptodds
Under review as a conference paper at ICLR 2027

CLASVS: Continuous-Latent Autoregression for Length-Changing Lyric Edits from a Reference Singing Performance

Abstract

Reference-conditioned lyric editing changes words within a sung phrase while preserving its melody, singer identity, and naturalness. When edits change the syllable count, target lyrics must override source-correlated cues while retaining the reference phrasing. We introduce CLASVS, a continuous-latent autoregressive editor with learned stopping. Its State–Control–Transition (SCT) history-access contract keeps target lyrics and reference melody available throughout generation, returns causal semantic feedback from generated latent patches to the planner, and confines direct predecessor acoustics to local synthesis. Progressive State–Control Grounding (PSCG) learns these routes through semantic pretraining, control perturbations, and staged reconstruction without paired lyric-edit recordings or score annotations. Across two Mandarin benchmarks, CLASVS improves deletion and insertion over continuous-NAR YingMusic+ and all four edit operations over discrete-AR Vevo2. Blinded paired listening favors its naturalness and target-lyric intelligibility, complementing objective measurements of melody and singer preservation. Matched routing and training comparisons support the access design and grounding recipe. Together, these results connect reliable lyric-count editing with the quality of the finished singing, using performances and transcripts for training. Audio demos are available at https://anonymous.4open.science/w/CLASVS-65B7/index.html.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.