HIM-CoT: Reasoning over the Human-Instrument-Music Chain for Musical Performance Deepfake Detection
Abstract
Audio-visual deepfake detection increasingly exploits cross-modal correspondence, yet existing methods predominantly target talking or singing heads, where detection relies on speech–lip synchronization and facial artifacts. However, this formulation breaks down for instrumental performances. Sound production is mediated by physical interaction with an instrument. Human motion drives instrument state, which in turn generates music. This structured causal chain imposes constraints that face-centric detectors cannot capture. Therefore, the core challenge is whether the observed performance provides a physically plausible explanation for the accompanying sound. We formalize this as an evidence-grounded reasoning problem: given structured measurements of music, motion, and performer–instrument contact, can a model determine which modality violates the causal production process? To address this question, we first construct HIM-DF, a benchmark with bidirectional forgeries that isolate audio and visual inconsistencies within the same Human–Instrument–Music structure. We then propose HIM-CoT, a reasoning framework that converts structured performance measurements into queries over pretrained multimodal representations. By explicitly representing measurement uncertainty, modality-specific evidence, and cross-modal temporal discrepancies, HIM-CoT grounds multistage chain-of-thought generation in forensic cues rather than end-to-end feature fusion. Experiments demonstrate that structured evidence retrieval and explicit discrepancy modeling improve both detection accuracy and interpretability over joint embedding baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.