Same Frames, Explicit Referents: Identity-Linked Evidence Bundles for Long-Video QA
Abstract
Long-video QA can retrieve every decisive frame yet still lose the identity relation that makes those frames jointly interpretable. To measure this post-retrieval failure directly, we introduce an exact-visual-parity test in which paired VLM inputs share every frame and crop byte, timestamp, item identifier, position, prompt, and visual order; only the answer-time organization changes from window-local records to a cross-window referent bundle. ReferentWeave realizes the tested bundle by composing typed temporal patches into persistent identity chains with relation and provenance fields, deterministic serialization, and uncertainty-triggered evidence expansion. On 3,842 Identity VideoQA queries, this intervention raises exact-match accuracy from 56.1% to 59.4% (+3.3 points; 95% CI [1.4, 4.9], video-cluster paired ) while serving the same 151 images per query. The exact-parity effect remains +2.5 points with InternVL2.5 and +2.3 under human-linked identities; when selection and decoding are restored under a common ceiling, ReferentWeave reaches 59.7% accuracy, +5.6 over Track-RAG and +8.3 over Dense-frame, while reducing identity swap from 7.8% to 6.0% and median served evidence from 186 to 148 images. Evidence access is therefore not a complete control for long-video QA: the referential structure attached to that evidence is an independently testable part of the model input.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.