Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection
Abstract
Speech generation methods are evolving rapidly, creating a moving target for audio deepfake detection (ADD). A deployed detector must incorporate newly emerging deepfake methods without forgetting previously learned real and deepfake knowledge. Continual learning provides a natural solution, but existing continual ADD evaluations commonly define tasks by dataset, coupling changes in real-speech sources with changes in deepfake mechanisms and obscuring what knowledge is being updated. We address task organization and detector adaptation jointly. First, we construct five protocols from identical training, development, and evaluation pools. Among them, the proposed real-anchored mechanism-incremental (RAMI) protocol reflects a practical detector-update scenario: real speech from known source domains is available, while newly arriving deepfake mechanisms must be learned continually without forgetting earlier knowledge. Second, we propose RF-Prompt (real–fake prompt learning), an asymmetric continual prompt-learning method. A shared real prompt provides protected adaptation capacity, while task-specific fake experts expand as new mechanisms arrive. Parameter-level cosine anchoring stabilizes the real prompt; each new fake expert inherits a selected historical expert and learns a residual with soft orthogonal regularization. Input-adaptive fusion integrates the accumulated fake experts without requiring task identity at inference. RF-Prompt obtains its lowest common-average and pooled EER under RAMI among the five tested protocols. On RAMI, it achieves 10.110% final average EER and 10.370% pooled EER, the lowest aggregate errors among the evaluated continual-learning baselines. Ablations assess the adaptation components, while limited-training-data and cross-backbone experiments examine their applicability across settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.