IterCap: A Self-Annotating Captioning Framework for Instruction-Following Text-to-Speech
Abstract
Instruction-following text-to-speech (instruct-TTS) learns acoustic controllability from captions, making their quality and coverage fundamental to what the model can learn: an attribute that a caption fails to specify or ground receives little supervision, regardless of the synthesizer's capacity. Existing methods either take captions as given, whether human-written, generated zero-shot by an external audio-language model, or produced by a handcrafted pipeline, or train a caption model on such captions once through supervised fine-tuning or a single round of reinforcement learning. In all cases, the supervision originatand is fixed before the model trains on it. We instead ask whether a caption model can annotate its own supervision, formulating instruct-TTS captioning as a self-ann an asymmetry: verifying the attributes asserted by a caption is easier than generating a correct open-ended caption. IterCap closes this loop through three opeenerated captions with lightweight acoustic-consistency checks and independent caption models; expansion adds examples of underrepresented delivery conditions; and optl correctness into a reference-anchored, verifiable objective through a multiple-choice exam and Group Sequence Policy Optimization (GSPO). On a 325-clip human12 controllable dimensions, IterCap raises macro-averaged per-dimension caption accuracy to 83.4%, compared with 75.3% for the strongest open-source caption modesfers to the downstream system: with the 300K-sample corpus and Qwen3-TTS backbone held fixed, IterCap captions raise instruction-following accuracy to 88.3%, compars's native captions and 82.9% using the zero-shot caption model from which IterCap is initialized.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.