Let Detectors Talk: Detector-Guided Data Selection for Deepfake Detection
Abstract
Deepfake detectors often struggle to generalize to unseen forgery patterns. We investigate whether adapting the use of available training data to a detector’s evolving weaknesses can improve generalization. To this end, we alternate detector training with the selection of samples for subsequent training stages. The difficulty is deciding what to select: although labeled samples reveal where the current detector struggles, similar prediction scores may conceal different failure patterns and offer limited guidance on how to address them. We use the samples the detector finds difficult as anchors for identifying its current weaknesses. To characterize these anchors beyond the detector’s own scores, we select auxiliary detectors based on their detection performance and response redundancy. Their responses provide multiple views of both the anchors and the samples available in the training pool. We then select additional training samples that exhibit similar response patterns to the anchors, retrain the target detector, and repeat the process with updated anchors. Experiments on six cross-domain test sets demonstrate the effectiveness of our method and show competitive performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.