acceptodds
Under review as a conference paper at ICLR 2027

From Decodability to Utility: Speech Deepfake Detection in Audio Language Models

Abstract

Audio language models have emerged as general-purpose speech interfaces, yet how their speech deepfake detection capability should be characterized remains unclear. Existing evaluations often rely on a single model endpoint, obscuring whether observed behavior reflects information supported by the underlying audio-language system or its effective use. We present Representation Accessibility Versus Exploitation by the Language Pathway (RAVEL), a controlled evaluation framework that separates native behavior, supervised decodability of frozen representations, and downstream exploitation through task adaptation, together with encoder-depth analysis. We evaluate thirteen Audio-LLMs across five datasets, with supervised comparisons on ten open-weight models and source-protocol and depth analyses. Native outputs are benchmark-dependent and often fail in decision validity or correctness, whereas frozen audio- and language-side representations frequently support strong supervised detection. This coexistence shows that weak native behavior need not imply the absence of accessible authenticity information. Native adaptation consistently improves in-domain detection over language-side frozen readout. Depth-wise adaptation further shows that intermediate representations can outperform final blocks in cross-domain discrimination, while encoder regions favored by frozen readouts can reverse under adaptation and shift across target domains. Under 19LA → AT-ADD transfer, Native-SFT outperforms both frozen readouts in EER for 9/10 models but in Macro-F1 for only 1/10, revealing metric-dependent transfer gains. Overall, supervised decodability and adaptation utility are distinct: stronger frozen-readout performance does not determine which representations are most effective for encoder-to-language adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.