acceptodds
Under review as a conference paper at ICLR 2027

Attention Meets Artifacts: Bi-level Temporal Re-Alignment for AIGV Detection

Abstract

Recent advances in video generation have raised growing security concerns, making AI-generated video (AIGV) detection increasingly essential. Existing methods predominantly rely on Multi-modal Large Language Models (MLLMs), but overlook the discrepancies among samples produced by different generation paradigms. However, we empirically find that MLLM-based detectors exhibit pronounced performance imbalance across different generation paradigms, with a notable weakness on image-to-video (I2V) samples. We attribute this to Temporal Artifact–Attention Misalignment. In particular, unlike entirely synthesized T2V samples, I2V samples are conditioned on a real initial image, and thus the forgery artifacts become progressively more pronounced toward later frames. The LLM decoder, however, attends predominantly to earlier frames owing to its causal attention mechanism; as a result, the most discriminative forgery cues in later frames are left largely unattended. To accommodate diverse generation paradigms, we propose a joint optimization strategy termed Bi-level Temporal Re-Alignment (BiTRA) that operates at both the extrinsic and intrinsic levels. Specifically, at the extrinsic level, we introduce ”Selective Hierarchical Artifact Injection”, an adaptive query mechanism that establishes an additional cross-attention pathway for global forgery information access, leveraging a temporal-artifact prior for discriminative information selection. At the intrinsic level, ”Temporal Attention Smoother” is guided by the Kullback-Leibler margin loss to mitigate the excessive attention bias inherent in the LLM decoder. Extensive experiments demonstrate the effectiveness of our designs on all generation paradigms and confirm our superiority over existing state-of-the-art methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.