Beyond Isolated Clips: Structured Learning for Temporal Forgery Localization
Abstract
Temporal forgery localization must identify the extent of a manipulated interval, not only assign suspicious scores to short clips. We study speech-assisted visual localization in talking-person video, where the input is one audio-video clip and the output is a set of visual manipulation intervals on its original timeline. Our structured localizer scores a common graph of 25-Hz candidate segments, trains the scores through a globally normalized distribution over complete temporal labelings, and uses an input-derived speech prior only at inference. This separates the learning objective from the choice of speech-guided readout. On a locked 212-video ArEnAV preview, the resulting U+S system reaches 46.59/33.76/2.65 AP at IoU 0.50/0.75/0.95 for Diff2Lip-trained checkpoints and 34.03/23.91/2.62 for LatentSync-trained checkpoints. A source-calibrated independent model given the same decoder reaches 28.42/15.39/0.10 and 26.87/6.55/0.08, respectively; the AP75 gains are 18.37 and 17.36 percentage points. Cross-generator direction rows and a strict visual LAV-DF transfer evaluation provide complementary evidence, while a matched U-max control exposes a segment-versus-frame ranking trade-off. Together, these results support globally normalized segment learning as a practical complement to independent clip scoring under the fixed protocol, with each evaluation scope reported separately.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.