Self-Rewarding GRPO for Instruction-Guided Zero-Shot Speech Synthesis
Abstract
Following natural-language instructions while preserving an unseen speaker's identity remains challenging for zero-shot text-to-speech. We present ACE-Instruct-TTS and Self-Rewarding Group Relative Policy Optimization (SR-GRPO), a post-training method that derives instruction rewards from the TTS model itself rather than an external audio understanding model. SR-GRPO uses a frozen supervised fine-tuning (SFT) checkpoint to score the same generated speech tokens with and without the instruction, holding the target text and speaker conditioning fixed. Inspired by conditional pointwise mutual information, the length-normalized log-likelihood ratio provides a continuous measure of relative compatibility with the instruction. We combine this signal with speaker-similarity and duration rewards. We also introduce ACE-Instruct-TTS-Eval-ZH, a 650-sample Chinese benchmark evaluating instruction following, speaker preservation, and intelligibility across six categories, including compositional and conflicting conditions. SR-GRPO improves instruction-following rate from 67.08% to 75.23%, the highest among evaluated systems, and increases mean human-rated instruction-following scores from 3.64 to 3.90. Reward ablations favor the proposed likelihood contrast over conditional-only, rank-based, and external binary rewards. Speaker similarity remains above all evaluated external baselines but decreases relative to SFT, while character error rate remains comparable. Beyond this benchmark, SR-GRPO improves the Chinese InstructTTSEval macro-average score from 79.2% to 81.2% and improves joint success across three SpeechEditBench tasks under an adapted re-synthesis protocol. On a second backbone, CosyVoice3, SR-GRPO improves instruction-following rate by 5.54 percentage points, with a similar speaker-preservation trade-off. We train instruction following only in Mandarin and leave multilingual instruction following to future work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.