ESIAN: Training-Free Continual Speaker Identity Erasure for Zero-Shot TTS
Abstract
Zero-shot text-to-speech (ZS-TTS) models can clone an arbitrary voice from only a few seconds of reference audio, even when the speaker was absent from the training data. When a speaker requests that their voice no longer be cloned, the provider therefore needs to address this request regardless of whether that speaker's recordings were used for training. We study erasing target-speaker identity during generation while preserving voice cloning for other speakers. Finetuning-based unlearning approaches update the model's weights for every removal request, a cost that becomes impractical as requests accumulate. Inference-time alternatives avoid weight updates but assume that the incoming prompt is already known to belong to a target speaker. As new requests continue to arrive, maintaining erasure for previously enrolled speakers poses an additional challenge. We propose ESIAN (Erasing Speaker Identity via Adaptive gatiNg), a training-free framework that treats an erasure request as a single enrollment step. From several minutes of enrollment recordings, ESIAN derives a closed-form eraser together with an intrinsic gate that determines, for every incoming prompt, how strongly the eraser should be applied. Our framework enables continual speaker identity erasure without updating the backbone weights or revisiting the utterances of previous requests. Experiments demonstrate that ESIAN achieves erasure comparable to finetuning-based unlearning while preserving voice-cloning quality for the remaining speakers, and maintains low average target-speaker similarity as the enrolled set grows to a hundred speakers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.