acceptodds
Under review as a conference paper at ICLR 2027

Learning Persistent and Editable Voice Profiles with Generative Retrieval

Abstract

Existing controllable text-to-speech systems can design voices or modify reference speakers, but these capabilities are often mediated by implicit voice representations distributed across conditioning and generation paths without exposing a persistent state that can be compiled, recovered or edited. In this paper, we introduce GRAPE-TTS, a framework for Generative Retrieval and Acoustic Profile Editing, which represents each voice as a persistent acoustic profile comprising a short sequence of residual-quantization identifiers. A supervised residual vector-quantization codec maps speaker embeddings into profiles, and an assembler converts each profile into acoustic conditioning for a TTS backbone. All voice operations are then acts of generative retrieval over this shared state. Extensive experiments demonstrate that GRAPE-TTS achieves strong performance across voice design, cloning, and editing, while preserving the naturalness and intelligibility of the backbone. These results establish explicit acoustic profiles as a unified and direct representation for controllable speech synthesis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.