AOI: All-in-One Audio Deepfake Detection via Evidence-Grounded Rationales and On-Policy Distillation
Abstract
Generative audio models can now synthesize speech, music, and environmental sounds that are difficult to distinguish from genuine recordings, yet existing audio deepfake detectors are typically trained for a single audio type and a narrow family of forgery algorithms, forcing practitioners to deploy a collection of type-specific models. We propose AOI, an all-in-one audio deepfake detection framework that covers multiple audio types and multiple deepfake types with a single AudioLLM. AOI first trains three audio-type-specific experts by supervised fine-tuning (SFT) on evidence-grounded chain-of-thought (CoT) targets constructed by a separate multimodal reasoning model. This annotator receives measured acoustic descriptors when writing the targets, but neither the experts nor the final student is given those measurements as input, so the detector must perceive the evidence it recites. The experts are then distilled into one student by on-policy distillation with reward extrapolation: the student generates its own CoT trajectories and the domain experts supply dense token-level supervision on them. We evaluate on speech, music, and environmental-audio benchmarks for fine-grained forgery classification and all-type generalization ( held-out clips). AOI classifies five-way over three audio types with a single checkpoint at weighted accuracy, within of three per-type label-only experts that require oracle routing, and recovers a substantial share of the unification cost by on-policy distillation; among all-in-one rationale methods, it achieves the highest five-way accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.