acceptodds
Under review as a conference paper at ICLR 2027

ZeroOpera: Low-Resource Peking Opera Singing Voice Synthesis with Zero-Shot Timbre Transfer

Abstract

Recent singing voice synthesis (SVS) systems have achieved strong performance, but Peking Opera synthesis remains difficult because it requires domain-specific pronunciation rules and has only limited annotated data. We propose ZeroOpera, a low-resource Peking Opera SVS framework that supports zero-shot timbre transfer from an acoustic prompt. Instead of fine-tuning the foundation model, ZeroOpera keeps it frozen and uses its prompt-conditioned mel-spectrogram as an acoustic prior. A trainable decoder reconstructs regions that require opera-specific pronunciations and enhances the full spectrogram. We further introduce Technique-driven Phonological Modulation (TPM). TPM converts lexicon-derived phonological category labels into continuous feature-level controls. By separating Standard Mandarin pinyin from opera-specific pronunciation categories, TPM can handle pinyin–category combinations that are not observed during adaptation. This design does not require an expanded pinyin vocabulary or manually aligned opera-specific phoneme annotations. For evaluation, we develop an automatic G2P pipeline and a Peking Opera ASR model that recognizes pinyin syllables. On target passages from a Peking Opera singer excluded from opera-domain adaptation, ZeroOpera reduces Phoneme Error Rate by 19.4% relative to the frozen foundation model. It also improves timbre-similarity MOS by 0.43 when using unseen utterances from singers included in adaptation, and by 0.30 when using cross-domain prompts from 50 pop singers excluded from opera-domain adaptation. These results show that ZeroOpera improves pronunciation and operatic expressiveness while preserving the timbre specified by the prompt under low-resource domain adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.