Learning Predictively Grounded Event Chains for Video Event Prediction
Abstract
Video event prediction requires inferring an unseen continuation from visual history. An explicit event chain may be visually supported but omit predictive evidence, or appear predictive by inventing events. We define predictive grounding as the ideal conjunction of interval-level visual consistency (VC) and candidate-conditioned predictive sufficiency (PS). We introduce \methodfull (), which optimizes verifier-relative proxies for both criteria without event-chain annotations or a task-specific supervised cold start. A shared policy first samples an option-blind timestamped history, then predicts from the video, candidates, and an immutable copy of that history. Pass 2 may revisit the video but cannot revise the chain. Frozen visual and blindfolded text verifiers supply training rewards; PS measures discriminative usefulness, not factual correctness. Five-state perception-aware credit routing assigns separate advantages to chain and prediction tokens. All verifiers are removed at inference. Reported point estimates on FutureBench, VEPBench, and AVEP favor the combined method under separate benchmark-specific training. Single-verifier ablations show an alignment–sufficiency trade-off, and chain interventions support behavioral dependence without establishing a causal mechanism. Aggregate metric inconsistencies and diagnostic aggregation remain to be reconciled with evaluation records, limiting the quantitative conclusions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.