acceptodds
Under review as a conference paper at ICLR 2027

Adaptation or Recalibration? Continual Test-Time Adaptation for Audio Deepfake Detection

Abstract

Audio deepfake detectors meet new attacks after deployment: a detector trained on today's synthesis systems must run against generators and recording conditions that are not known in advance. We study this as source-free, label-free continual test-time adaptation (CTTA) and build a controlled, strictly prequential benchmark for audio deepfake detection, in which every utterance is scored before it can affect an update. The benchmark has five streams and 266,400 scored positions, with source training on ASVspoof 2019 LA and target streams pairing MLAAD spoofed speech with M-AILABS bona fide speech. It separates two failures that a single accuracy number hides: discrimination fails when the score ordering gets worse, and the operating point fails when the scores move relative to a threshold fixed before deployment. Five prediction-driven CTTA methods, evaluated on XLS-R and WavLM, move the operating point far more than they repair discrimination. On the stream of 52 sequential generators, Tent lowers per-block minimum detection cost (minDCF) by 0.010 and 0.013 on the two held-out checkpoints, against 0.100–0.112 for a labeled update of the same parameters. Recalibrating the frozen detector removes 22–77% of its calibration loss, depending on the checkpoint, without changing the ordering. An utterance's own score cannot say which of two utterances is ranked wrongly; relations between utterances can. We introduce CAST (Consensus-Anchored Score Transfer), which smooths frozen scores over reciprocal nearest neighbors and distills the resulting order into the adapting detector with a pairwise ranking loss. On the two held-out checkpoints of that stream, CAST lowers minDCF by 0.065 adapting only the LayerNorm affine parameters and by 0.103 adapting the full transformer encoder, where EER reaches 11.3%, against 0.011 for the strongest baseline. The gain holds on all three controlled target streams and both backbones. Smoothed scores improve discrimination even with no parameter update, while CAST's loss reads only score differences and leaves the operating point unresolved. Label-free continual detection therefore needs to solve two problems, and repairing discrimination solves only one.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.