Neuro-Symbolic Forensic Reasoning for Open-Generator AI-Generated Video Detection
Abstract
AI-generated video detection is critical for information integrity, yet remains hard in open-generator deployment: new generators shift artifact distributions, forensic cues are sparse and heterogeneous, and benchmark shortcuts can overstate progress. We argue that the root cause is the prevailing binary classification formulation f(V ) → 0, 1, which learns generator-specific statistical shortcuts rather than transferable forensic principles, making detectors brittle to unseen generators and costly to retrain. We recast the problem as forensic reasoning: instead of a binary label, the system must produce a verdict together with a grounded evidence chain, f(V ) → (ˆy, E). To instantiate this formulation, we propose FOVEA, a neurosymbolic framework for Forensic Open-generator Video Evidence Acquisition that trains no detector weights. A VLM planner forms forensic hypotheses and routes videos to specialized analysts, whose judgments are grounded by deterministic CV measurements including optical flow, frequency analysis, error-level analysis, temporal tracking, and geometric consistency. A reliability-aware judge fuses heterogeneous evidence via strong-evidence override, confidence-filtered voting, and disagreement-penalized aggregation to produce an interpretable verdict. On GenVideo, across ten unseen generators, FOVEA with Claude Opus 4.6 achieves 97.39% accuracy, outperforming the strongest supervised baseline, NSG-VD, by 5.93 points and reaching 0.99 AUROC, using only a one-time 100-video threshold calibration. The advantage persists after removing visible watermarks, on clips of at least 4 s, and on the GenVidBench benchmark.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.