GenTell-Gaze: A Dataset and Benchmark for Localising Physical Anomalies in AI-Generated Video with Human Gaze
Abstract
Generated AI videos are becoming increasingly realistic, making the identification of when and how they diverge from real-world behaviour not only harder, but also more important. An AI video failure that is still relatively common is incorrectly representing physical laws and dynamics, leading to anomalous artefacts in the videos. Correctly identifying and locating these can help create explainable and more trustworthy fake-video detectors, as well as help evaluate vision-language models’ ability to understand physics. We introduce the GenTell-Gaze dataset: 380 generated videos spanning 80 physics-based scenes and six video generators, with 1,070 manually annotated spatio-temporal anomaly `tubes', alongside eye-tracking data from 129 participants. We develop GenTell-Bench, a benchmark for evaluating vision-language model (VLM) reasoning capabilities. This includes classifying videos as real or AI-generated, as well as recognising, classifying, and localising any present physical anomalies. Using the eye-tracking data, we are able to evaluate whether human visual attention can improve VLM video reasoning and anomaly localisation. Across all models tested, performance falls substantially as the tasks progress from video classification to temporal and spatio-temporal localisation, revealing a persistent gap between recognising generated content and identifying the responsible anomalies with physical reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.