acceptodds
Under review as a conference paper at ICLR 2027

Learning Along Depth: Depth-Weighted Layer Supervision for Speech Anti-Spoofing

Abstract

Advances in neural speech synthesis have made audio deepfakes increasingly hard to detect, and self-supervised Speech Foundation Models (SFMs) have become the standard front-end for anti-spoofing. Standard fine-tuning, however, supervises only the final representation. Intermediate layers therefore remain weakly adapted to the task, and the burden of discrimination shifts to ever more expressive back-end classifiers. In this work, we ask how supervision can instead reshape the internal representation hierarchy of the SFM encoder. Our layer-wise analysis of fine-tuned SFM shows that anti-spoofing discriminability emerges late in the network, and later still as model size grows, so most intermediate layers remain poorly aligned with the task. Motivated by this, we propose Depth-Weighted Layer Supervision (DWLS), an effective multi-layer supervision framework that progressively aligns representations across the encoder depth with the anti-spoofing objective. Across diverse deepfake benchmarks, DWLS paired with a lightweight classifier achieves competitive in-domain performance and strong out-of-domain generalization, suggesting that improving internal feature alignment can be a more efficient route to adaptation than increasing back-end complexity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.