Learn What to Say: Generalizable Active Enrollment for Few-Shot Voice Personalization
Abstract
Few-shot voice personalization can reproduce a target speaker from only a few seconds of reference audio, yet users receive little guidance about what they should say. We formulate active voice enrollment as a budget-constrained sequential selection problem and introduce LEAP (Language-Transferable Enrollment via Active Policy Learning). LEAP predicts the marginal utility of each candidate utterance conditioned on the content already selected, enabling it to construct complementary enrollment sets without synthesizing speech or measuring downstream quality at inference time. We evaluate LEAP across Korean, Japanese, and Vietnamese, multiple voice-personalization models, matched recording budgets, and human listening studies. LEAP improves automatic personalization utility over random selection, natural user-selected enrollment, phoneme coverage, language-specific handcrafted objectives, and a learned static-ranking baseline. Crucially, when trained only on outcomes from two source languages, LEAP outperforms a handcrafted objective designed specifically for the held-out target language; its gains also persist for an unseen voice-personalization model. Human listeners prefer outputs generated from actively selected enrollment for speaker similarity, naturalness, and content clarity. These results show that enrollment utility contains transferable structure: when only a few seconds are available, performance depends not only on how much a user speaks, but also on what the system asks them to say. The selection summary and lexicon file for LEAP are available at https://anonymous.4open.science/r/Learn-What-to-Say-F86C.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.