acceptodds
Under review as a conference paper at ICLR 2027

Edit Who Speaks, Control How They Speak: Global Timbre Editing and Local Instruction Control for TTS

Abstract

Instruction-based text-to-speech (TTS) mainly focuses on cloning persistent speaker identity or controlling local speech expression. Voice cloning preserves the speaker identity from a reference voice, whereas text-based voice design generates a voice from a natural-language description. Neither provides a mechanism for explicitly editing the timbre of a given reference voice. Meanwhile, utterance-level instructions control speech expressiveness, but cannot specify how it should vary across segments. We introduce EDICT, a framework that decouples persistent speaker identity from transient expressive instructions while coupling them through an editable acoustic representation. EDICT first maps a reference voice and structured timbre edits to an edited codec-token representation. This representation is shared across all text segments to anchor a consistent target voice. For each segment, a separate natural-language instruction is injected into a frozen TTS backbone to control local expression. To enable instruction switching without disrupting acoustic continuity, EDICT rebuilds the KV cache at each segment boundary. It retains a bounded acoustic context to preserve local speech history while refreshing the instruction condition. This design provides independent control over what the voice sounds like and how it speaks across segments. Evaluations on our proposed TimbreEdit-Bench and IntraTTS-Bench demonstrate improved timbre-editing accuracy and segment-level instruction adherence over baseline methods, while maintaining cross-segment voice consistency and speech quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.