acceptodds
Under review as a conference paper at ICLR 2027

LongCoref-AV: Benchmarking Audio-Visual Identity Coreference in Long-Form Video

Abstract

An unresolved challenge for multimodal models processing long-form video is determining which faces, voices, and actions across distant moments belong to the same person. To measure this capability, we introduce LongCoref-AV, a benchmark of 1,126 human-verified multi-choice questions across 678 long-form videos. Questions identify a target individual only through visible or audible cues, such as clothing, gestures, or laughter, and are designed to rely on evidence from moments minutes apart. Each question is labeled with the category of evidence it requires to be answered correctly: voice, visual, or joint. Across 16 systems, closed frontier models far exceed a text-only baseline, open models lag far behind, and voice-dependent questions prove to be the most challenging across all models. When given only the annotated video segments containing the relevant evidence rather than the full video, open models improve far more than frontier models, suggesting that existing open models mainly fail to find evidence, while frontier models are more likely to locate it but often attribute it to the wrong person. In a separate identity-matching test, the two strongest open omni models both incorrectly label more than half of pairs depicting different individuals as the same person, despite their strong performance on multi-choice questions. This discrepancy shows that end-to-end question accuracy does not certify the intermediate identity links: models can select correct options from partial contextual cues without reliably distinguishing individuals across moments. Our findings expose persistent identity attribution as a critical weakness of multimodal LLMs, positioning LongCoref-AV to support more reliable attribution of actions to individuals in long-form video. Code and data will be made publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.