What Happened vs. What Can Be Known: Ontic–Epistemic Video Revisit Reasoning
Abstract
Video-language models are increasingly capable of recognizing temporal state changes, yet they often conflate what happened in the world with what can be concluded from the visual evidence they actually observed. We study this distinction through ontic–epistemic reasoning across video revisits, where a model must compare entity states across temporally separated visits, recognize a physical transition, and respond appropriately when visit-specific evidence is incomplete. We construct structured Revisit Packs from paired observations of the same environment. Each pack contains controlled variants that remove or perturb visit-specific evidence while preserving the underlying physical case. We first extend Behavior Pack Optimization (BPO), originally designed around counterfactual views of a single video-question instance, to these paired revisits so that cross-visit transition behavior can be optimized jointly. Although BPO provides anchor-relative updates across the pack, it does not explicitly prioritize the intervention relations that remain unsatisfied. We therefore introduce Intervention-Structured Bottleneck Behavior Pack Optimization (BBPO), which uses detached pack-local relation-success scores to redistribute BPO advantage toward poorly satisfied revisit relations while preserving the mean optimization scale of the pack. Experiments show that BBPO raises mean all-view pack correctness from 63.6% to 66.3%, including improvements from 54.0% to 57.5% on Revisit Packs and from 68.4% to 70.7% on Visit Packs. Extensive results suggest that reliable reasoning across video revisits benefits not only from jointly optimizing evidence interventions, but also from allocating post-training credit according to the behavioral relation that most limits all-view revisit correctness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.