Beyond Pixels: Bridging the Cognitive Gap in AI-Generated Video Detection
Abstract
Recent advances in video generation have produced highly realistic and temporally coherent videos that are increasingly difficult to distinguish from real content. While beneficial for many applications, such progress also raises concerns such as misinformation. Existing detection methods mainly rely on pixel level artifacts and shallow spatiotemporal inconsistencies, which often fail as generation quality improves. We argue that detecting AI-generated videos of high quality requires deeper cognitive understanding beyond low level signals. Thus we propose a framework that defines five cognitive scenarios, each containing different levels of cognitive information, ranging from perception and physical plausibility to semantic, logical, and social reasoning. Through a study on how human and current MLLM detect AI-generated videos, we show that detection performance differences across various cognitive scenarios, which are related to the quality of the AI-generated video and the video understanding abilities of humans and models. Building on these findings and our proposed cognitive framework, we introduce an Iterative Cognitive Layered Instruction Evolution framework that enables multimodal models to evolve detection instructions under structured cognitive guidance. Experiments demonstrate improved detection performance and strong generalization across models and unseen AI-generated videos.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.