STELA: Stable Temporal Evidence Learning for Audio Forgery Detection and Localization
Abstract
Audio forgery detection and localization aims to identify spoofed frames in manipulated speech and determine the forgery boundaries. Existing methods have primarily improved frame discrimination through stronger feature learning and boundary-aware modeling. However, local acoustic fluctuations within manipulated regions can still compromise prediction, leading to fragmented and discontinuous localization. We propose STELA, a Stable Temporal Evidence Learning framework for Audio forgery detection and localization to address this challenge. Specifically, for stable evidence learning, a segment-oriented acoustic augmentation and Paired State Modeling (PSM) provide complementary constraints against local acoustic variability. The augmentation introduces segment-level perturbations to discourage the detector from conflating incidental acoustic variations with manipulation boundaries. PSM further regularizes the representations within each utterance, encouraging acoustically varied frames sharing the same authenticity state to remain consistent. By mitigating representation drift and enhancing temporal coherence, these complementary constraints provide a more reliable basis for subsequent manipulation localization. Building on this improved evidence, Structured Manipulation Localization (SML) extends threshold-based proposal generation with structured scoring of alternative interval hypotheses, using interior, boundary, and contextual evidence to recover coherent manipulation intervals from fragmented predictions. Extensive experiments on PartialEdit, PartialSpoof, HAD, and LAV-DF demonstrate that STELA achieves state-of-the-art performance in both forgery frame detection and temporal localization, with strong generalization to unseen conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.