Events, Not Frames: Detecting Physical Violations from a Single Video
Abstract
Infants look longer when a box hangs in mid-air instead of falling; detecting such violations of physical expectation is a hallmark of intuitive physics and a long-standing goal for machines. Current methods often use a matched plausible clip to interpret a prediction error, or annotated violations to train an evaluator. Infants need neither to judge a single event from ordinary experience alone. Here we show that event structure may help close this gap by separating the physical expectations that frame-level prediction errors conflate. To this end, our model, Physics Using Latent State-Events (PULSE), parses tracked 3D object states into events and asks, stage by stage, whether an object should be visible (Visibility), whether the interaction begins as it should (Onset), and whether the state it leads to is consistent (Outcome). Each of these checks asks the same question of every video, and plausible videos alone define the normal range against which an unmatched clip is judged. On three benchmarks, including PhysEvent, our new benchmark with stage-specific violations and ground-truth 3D states, PULSE achieves the strongest overall single-video detection and paired ranking among the evaluated continuous scorers. Each decision also names the event, check, and objects behind it. More broadly, this work points toward machines that turn experience with ordinary events into an understanding of what can and cannot happen.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.