From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection
Abstract
One score per utterance does not show which component measurements supported a speech deepfake decision. We ask whether adding two absolute score differences to cross-fitted logistic fusion lowers equal error rate (EER): one difference compares the passive and watermark scores, and the other compares their average with retrieval. The decision record retains four primary scores and an auxiliary neighbor-distance field for post hoc diagnosis. The record stores a passive detector score on original audio, a watermark score on a marked copy, a retrieval vote from ten nearest support examples, and a speaker-profile margin based on nearest bona fide and spoof distances. It also stores distance to the nearest support example. For spoof examples, retrieval excludes the target synthesis family. On 4,080 ASVspoof 5 Track 1 development examples with all four primary scores, the absolute-difference model reaches 8.43% EER. Cross-fitted scalar fusion reaches 11.91%, and the best of 22 passive WavLM runs reaches 6.71%. Across eight held-out families, the mean fold EER reduction is concentrated in A11 and A12. The controls do not identify what causes the reduction. The fixed rule averages passive and watermark scores, then averages that result with retrieval. Its minimum detection cost is lower than the learned model's. In an in-sample, post hoc analysis of the same examples, sorting by distance to the final-score threshold places 88.66% of observed errors among the 33.75% of examples closest to the threshold. This analysis does not evaluate a deployment policy or analyst utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.