Frames Don't Add Up: A Factorial Causal Audit of Contextual Evidence Utility in Video-Language Models
Abstract
Video question answering systems often select frames by scoring each independently for relevance, change, or diversity. Yet the answerer sees a set, not a ranking. Does a frame retain its value beside the other selected frames? We test this assumption through controlled downstream interventions. Our "Factorial Evidence Audit" holds the frame budget, surrounding frames, candidate slots, pixels, prompt, and answer scoring fixed while measuring individual and joint effects across roughly one million calls, six open video-language models, and three evidence-grounded benchmarks. Annotation-defined frames are causally useful, exceeding matched outside-evidence frames in 17 of 18 model–benchmark settings, with clustered intervals excluding zero in 16. But their values do not add independently. Evidence pairs interact more strongly than matched outside-evidence pairs in every setting (median absolute-interaction gap gold-vs-rest log odds), even after adjusting for the scale of their individual effects. This gap remains positive under chronological re-sorting and controls for candidate order and slot, score scale, context, frame budget, distractor construction, and RGB decoding. Deterministic scoring makes zero the exact additive prediction, and no-op controls return exactly zero. On LVBench and HERBench, common pointwise proxies poorly track measured utility; third- and fourth-order terms carry 23.5–27.7% of the absolute nonconstant M\"obius mass. A four-benchmark selector census exposes the practical consequence: in none of its 18 grounded settings does the method retrieving the most annotated evidence help the answerer most. Evidence matters. But fixed-budget video evidence selection is a set problem, not a ranking problem.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.