acceptodds
Under review as a conference paper at ICLR 2027

Self-Rewarding GRPO for Instruction-Guided Zero-Shot Speech Synthesis

Abstract

Following natural-language instructions while preserving an unseen speaker's identity remains challenging for zero-shot text-to-speech. We present ACE-Instruct-TTS and Self-Rewarding Group Relative Policy Optimization (SR-GRPO), a post-training method that derives instruction rewards from the TTS model itself rather than an external audio understanding model. SR-GRPO uses a frozen supervised fine-tuning (SFT) checkpoint to score the same generated speech tokens with and without the instruction, holding the target text and speaker conditioning fixed. Inspired by conditional pointwise mutual information, the length-normalized log-likelihood ratio provides a continuous measure of relative compatibility with the instruction. We combine this signal with speaker-similarity and duration rewards. We also introduce ACE-Instruct-TTS-Eval-ZH, a 650-sample Chinese benchmark evaluating instruction following, speaker preservation, and intelligibility across six categories, including compositional and conflicting conditions. SR-GRPO improves instruction-following rate from 67.08% to 75.23%, the highest among evaluated systems, and increases mean human-rated instruction-following scores from 3.64 to 3.90. Reward ablations favor the proposed likelihood contrast over conditional-only, rank-based, and external binary rewards. Speaker similarity remains above all evaluated external baselines but decreases relative to SFT, while character error rate remains comparable. Beyond this benchmark, SR-GRPO improves the Chinese InstructTTSEval macro-average score from 79.2% to 81.2% and improves joint success across three SpeechEditBench tasks under an adapted re-synthesis protocol. On a second backbone, CosyVoice3, SR-GRPO improves instruction-following rate by 5.54 percentage points, with a similar speaker-preservation trade-off. We train instruction following only in Mandarin and leave multilingual instruction following to future work.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.