MoLN: Mixture of Layer Normalization Experts for Speech Deepfake Detection
Abstract
Speech deepfake detection (SDD) requires strong generalization ability due to the diverse and heterogeneous factors affecting speech generation, transmission, and recording conditions. Recent SDD systems increasingly leverage large pre-trained speech encoders and parameter-efficient fine-tuning (PEFT) to adapt transferable speech representations, with LoRA-based approaches becoming a common adaptation strategy. In this work, we investigate LayerNorm (LN) tuning as an alternative adaptation interface for generalizable SDD. Different from LoRA, LN tuning performs compact channel-wise feature recalibration, providing a favorable trade-off between parameter efficiency and generalization while better preserving pre-trained representations. We further show that a single shared LN is insufficient to model the heterogeneous and overlapping conditions in SDD. To address this limitation, we propose Mixture of LayerNorm Experts (MoLN), a parameter-efficient framework that extends LN tuning from a single affine parameter to multiple lightweight LN experts. MoLN employs a frame-level soft router to dynamically combine expert-specific scale and shift parameters, enabling adaptive feature recalibration without requiring explicit domain labels or additional task-specific modules. Extensive experiments demonstrate that MoLN consistently improves cross-dataset generalization over state-of-the-art baselines with fewer trainable parameters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.