PERSISTQA: DIAGNOSING IDENTITY ASSOCIATION IN LONG-VIDEO QUESTION ANSWERING
Abstract
Long videos often contain people who reappear across different scenes and moments. Answering a question about one person therefore requires more than finding a relevant event. A model must recognize the intended person across time and associate the correct observations with that person. However, long-video question answering is typically evaluated using aggregate answer accuracy, which does not reveal where this process fails. We introduce PersistQA, a diagnostic benchmark for studying identity grounding in long-video question answering. PersistQA contains 3,299 eight-option questions over 409 videos. Each question identifies a person through visual or relational cues and asks about the same individual at another moment. Across 27 models, the best open-weight and closed-source systems achieve only 27.2% and 29.3% accuracy, compared with 60% for a human no-look-back baseline. Our diagnostics reveal several sources of error. Frozen CLIP features distinguish same-person from different-person pairs with 85.2% accuracy, while VLMs explicitly judging identity reach only 50.6-62.6%. Eight uniformly sampled frames miss approximately 53% of candidate-answer windows, and replacing one frame with candidate evidence improves accuracy by 4.3-9.7 percentage points. Even when candidate evidence is directly provided, 61-78% of questions remain incorrect. These results highlight identity association, evidence access, and evidence use as distinct challenges for long-video question answering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.