VOICE OR LENGTH? COUNTERFACTUAL AUDITING OF TIMING SHORTCUTS IN AUDIO DEEPFAKE DETECTION
Abstract
A benchmark can remove a shortcut from its evaluation data while leaving it in its training data. A detector may then learn a highly predictive nuisance cue that the benchmark neither rewards nor exposes. We study this gap for timing in audio deepfake detection, and separate a shortcut’s availability in a corpus from a detector’s reliance on it. Availability: ASVspoof 5 equalises timing across classes, by design, only in the partitions it tests on; on held-out speakers of its training partition, timing alone separates the classes at AUC 95.6, against chance on its evaluation set. Across four corpora the predictive component differs, and on ASVspoof 2019 LA (LA19) the components that survive boundary trimming, the standard control, retain AUC 87.7 of the 88.3 available. Reliance: matched-length evaluation confounds speech amount with boundary context and a shifting population, so we intervene instead, holding the speech byte identical. Planting real bonafide boundary silence on spoofed ASVspoof 5 clips raises a released XLS-R Conformer from 5.8 to 9.4% EER, against 7.1% when every clip receives it. Incentive: fine-tuning that checkpoint further on its natural view lowers its LA19 benchmark EER from 0.35 to 0.17% while raising its length-matched In-the-Wild EER by 4.7 points, so, in this setup, selecting on the benchmark picks the detector that relies more on the shortcut. Randomly placed training crops reverse this and remove the planted-silence reliance, at a cost in LA19 EER, and turning the intervention into an objective — a consistency loss between each crop and the same speech with label-independent silence added — gives the lowest length-matched In-the-Wild EER of any LA19-trained detector (6.6%, from 9.2%); a duration-aware architecture, tested over twenty seeds, does not help. For timing, removing the cue from a test set does not test for it; intervening on it does. We release the code, ports of five released checkpoints, and per-clip scores for every condition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.