Forgetting Who, Preserving What: Speaker-Identity Unlearning in Audio-Language Models
Abstract
Audio-language models jointly encode who is speaking and what is said, creating a practical deletion request: remove a target speaker-identity association from a deployed model while preserving the linguistic value of the same recordings. Existing sequence-level unlearning objectives do not distinguish identity-bearing acoustic evidence from content-bearing evidence and can therefore trade stronger forgetting for broad representation damage. To address this challenge, we introduce CAUL, a selective model-side unlearning framework for AudioLLMs. It consists of three stages: (i) Intervention-Aligned Support Localization, which identifies acoustic channels whose neutralization suppresses target-identity evidence more than linguistic content; (ii) Cross-Utterance Evidence Aggregation, which aggregates channel-level evidence across utterances to obtain stable speaker-specific support; and (iii) Sparse Local Association Editing, which neutralizes the selected acoustic support toward a retain-derived reference and jointly optimizes target-identity unlikelihood with retain KL to preserve non-target behavior. Alongside CAUL, we establish VoxSU-Bench, a speaker-unlearning benchmark built from VoxCeleb datasets. Experimental results on VoxSU-Bench show that CAUL achieves strong speaker-identity forgetting while preserving high speech-content retention. To the best of our knowledge, this is the first study of speaker-identity association unlearning for AudioLLMs. We will release the code and benchmark in the near future.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.