RePlan-VC: Target-Side Semantic and Prosodic Replanning for Controllable Voice Conversion
Abstract
Controllable voice conversion aims to modify speaker identity and expressive attributes while preserving linguistic content. Existing methods either retain source-derived content anchors while replanning selected factors or generate target-side states without explicitly modeling dependencies among functionally distinct states, limiting coordinated multi-state replanning. We propose RePlan-VC, a target-side semantic and prosodic replanning framework that regenerates semantic and prosodic trajectories under target speaker and emotion conditions. RePlan-VC employs a hierarchical autoregressive planner that generates semantic states before prosodic states, explicitly modeling their generative dependency, while target-condition modulation directly guides state prediction. To preserve linguistic content without enforcing source state reproduction, a frozen semantic to text mapper provides sequence level supervision over the replanned semantic trajectory. Experiments on timbre, emotion, and joint conversion demonstrate effective target attribute adaptation while preserving linguistic content. Under joint conversion, RePlan-VC achieves 7.68% WER, 0.5046 SECS, 0.6070 EmoSim, and 52.5% EmoAcc.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.