acceptodds
Under review as a conference paper at ICLR 2027

Training Data Redirects Evidence Routes in Audio-Visual Deepfake Detection

Abstract

Audio-visual deepfake detectors reach near-perfect in-domain AUC yet fail unpredictably under distribution shift, and accuracy alone does not say why. We argue that what matters is the detector's evidence route, the channel through which predictive evidence reaches its decision, and that the route is a property of the trained instance rather than of the architecture alone. AVRouteBench makes the route measurable in two steps: an intervention-based signature (leading silence, speaker identity, modality skew, generator overfit) that narrows the route to a few candidates, and a direct confirmation intervention, such as breaking the audio-visual pairing while leaving each stream intact, that tests each candidate. Two findings follow. First, training data can redirect the route within what a fixed architecture allows. With one frozen AV-HuBERT backbone and one logistic head, changing only the training mixture moves fake-video/real-audio AUC from 0.495 to 0.980. The effect survives a matched sweep that fixes class and cell balance and the within-identity real-fake pairing, and a same-source intervention traces part of it to the availability of an audio-side label cue. A published detector with identical architecture and code learns a visual route from one training source and an audio-dominated route from another, and on two other frozen backbones the size of the effect tracks the visual evidence a linear head can read from the representation. Second, the route sets the failure mode: which edits a detector still catches under shift depends on whether the evidence it reads survives the target's manipulation type. Audio-visual correspondence is the case we characterize in detail. Across 19 detector instances from ten families, AVRouteBench shifts detector evaluation from how accurate is it? to which evidence does it use, and will that evidence hold under shift?

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.