When Later Evidence Changes the Verdict: Snapshot-Bound Citation Evaluation for Streaming Multimedia QA
Abstract
An answer issued while OCR and crop extraction are incomplete can later receive a better citation score even though its text and returned citation handles never change. Completed-index evaluation therefore lets future evidence alter the recorded quality of a past multimedia QA response. Snapshot-bound evaluation exposes this temporal leakage by committing each output to source-native evidence roots and its served snapshot, then replaying those fixed handles through an explicit provenance relation under served and settled revisions. On 72 streams spanning 187 hours and 14,382 verified QA pairs, settled replay raises citation correctness by 1.8 points for ServeTrace, 3.3 for Rolling-ANN, and 4.4 for Tier-less indexing; all stream-block 95% confidence intervals exclude zero. The service mechanism is visible across backlog strata: sketch-only answers rise from 8.3% to 32.4%, served citation correctness falls from 82.4% to 76.1%, and QA F1 on answered queries falls from 72.0 to 69.1. At the prespecified 80% line, 4.1–6.7% of answer-level labels change, with fail-to-pass transitions outnumbering reversals by 11–14 times. The effect transfers to the released Video-RAG pipeline on TVQA+ (+3.4 points, 95% CI [2.4, 4.5]) and to InternVL3-38B. ServeTrace realizes the protocol at 82.1% served citation correctness, 2.1 GB/h median storage, and 2.58 s p95 latency, while remaining within 0.8 QA-F1 points of Store-all. The resulting audit distinguishes evidence verifiable at response time from support that becomes resolvable only after the decision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.