acceptodds
Under review as a conference paper at ICLR 2027

SRAM: Learning Spectral Relation-Aware Representations for Environmental Sound Deepfake Detection

Abstract

Environmental sound deepfake detection (ESDD) models are critical for countering fabricated environmental audio, which can mislead users of audio-based assistive technologies and trigger unintended smart-home responses. A reliable ESDD method requires accurate detection even when the source datasets and generators are unseen in training. Existing ESDD methods broadly follow two paradigms: task-specific discriminative models and generative methods based on audio large language models (ALLMs). We find that both approaches have significant limitations: (i) discriminative methods are highly dependent on the seen domain conditions of the training data, and their classification performance drops significantly when the test audio comes from unseen domains; (ii) generative methods generally underperform and incur higher per-sample inference latency and deployment costs. We conduct a large-scale evaluation of 15 ALLMs, spanning five model families, on the ESDD task. The results show that their zero-shot performance is close to random guessing, and unlike the gains reported in speech deepfake detection, their performance remains significantly inferior to that of strong task-specific discriminative models even after fine-tuning. Further analysis shows that a classifier built on an ALLM's audio encoder can outperform the full generative model, while local spectro-temporal residuals and time-frequency geometry reveal differences between real and AI-generated audio. Motivated by these findings, we propose a novel detection framework, called Spectral-Relational Artifact Modeling (SRAM). SRAM combines audio representation learning with spectral relation modeling. It builds on the Qwen3-Omni audio encoder, using cross-layer trajectory refinement to integrate information from multiple encoder depths and sparse evidence pooling to focus on informative temporal tokens. To complement these audio representations, we use a spectrogram-residual branch to capture local spectral differences and a structure-tensor branch to describe local time-frequency geometry. Extensive experiments demonstrate SRAM’s strong detection performance, generalization to unseen domains, and robustness to unseen perturbations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.