Achilles' Heel: Unveiling temporal artifacts in AI-Generated videos via video diffusion models
Abstract
Despite rapid advances in visual quality, current video generation models still produce temporal artifacts, such as appearance drift or unexplained object emergence or disappearance that may be imperceptible from individual frames but obvious as time unfolds. Existing benchmarks predominantly lack attention to this type of artifacts and provide coarse-grained annotation. We propose Temporal Artifact Bench (TA-Bench), a benchmark dedicated to fine-grained diagnosis of temporal artifacts in AI-generated videos, comprising 9 generative models, 6,945 video with 8,618 continuous temporal masks and 214.3K frame-level annotations. To construct TA-Bench at scale, our pipeline combines automatic triage, mask generation, controllable synthesis, and 3 layer human verification. We evaluate various model families for temporal artifact localization, including general purpose multimodal language models (MLLMs), AIGC detection MLLMs and video segmentation models. Results show VLM-based methods depicting limited sensitivity to temporal artifacts, whereas generative video models achieve the strongest localization performance. Motivated by this finding, we propose Video Diffusion Diagnoser (VDD), a generation-for-detection approach that repurposes a pretrained video diffusion model to predict temporal continuous masks directly from video latent in a single feed-forward pass. Experiments on TA-Bench show that VDD outperform MLLM-based or segmentation-based models, suggesting that generative video priors provide a strong foundation for fine-grained temporal artifact perception.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.