acceptodds
Under review as a conference paper at ICLR 2027

VIOLATE: Evaluating Video-Language Models on Fine-Grained Physical Implausibility

Abstract

Model-generated videos can look realistic while containing brief implausible events. Video-level evaluations cannot show which event occurred, when it occurred, or whether several events occur in the same video. We introduce VIOLATE, a benchmark for fine-grained physical implausibility detection. It contains 556 human-annotated synthetic videos with 1,552 events. Each annotation describes what goes wrong and when it happens. We also use 150 real videos to measure false detections. The benchmark supports two tasks. In event verification, a model decides whether a proposed implausible event is actually visible. In open-ended event detection, the model receives only the video and must find, describe, and localize every implausible event, or return no event when the video is plausible. Current models struggle with both tasks. GPT-5.5 reaches the best off-the-shelf Matthews correlation coefficient of for verification, while the best fine-tuned model reaches ; both values show weak discrimination. For open-ended detection, Claude OpusĀ 4.8 reaches F1 for matching event descriptions and F1 for their time intervals. We propose three-stage training with teacher distillation, reinforcement learning (RL), and self-distillation. These stages first learn from teacher-generated reasoning, then improve event recovery, and finally learn from the RL model's own predictions together with videos that contain no annotated implausible event. Its best variants reach semantic F1 and temporal F1. We will release VIOLATE to support evaluation and training for physical implausibility in model-generated videos.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.