acceptodds
Under review as a conference paper at ICLR 2027

AnchorGroup: Object-Anchored Video Understanding Benchmark with an Agentic Annotation Loop

Abstract

Video Question Answering (VideoQA) has seen advances in both Multimodal Large Language Models (MLLMs) and evaluation benchmarks. Existing benchmarks typically evaluate independent QA pairs and report performance for each task. Questions from different tasks may concern different objects, even within the same video. Separate task scores therefore provide limited insight into whether models succeed across tasks involving the same object. To address this gap, we introduce AnchorGroup, a video understanding benchmark organized into object-anchored QA groups. Each QA group contains perception, reasoning, and summarization questions about the same object. The shared anchor provides a common reference for comparing task performance and examining associations between task outcomes. Existing annotation pipelines primarily focus on individual QA pairs rather than shared-anchor constraints across tasks. We therefore introduce A3Loop, an anchor-guided agentic annotation loop that constructs QA groups with direct access to video clips. The loop generates QA pairs from summarization to reasoning and then perception. During construction, verifier feedback guides generators in revising QA pairs to meet task-specific and group-level requirements. Revised QA pairs are then verified again. Evaluations on AnchorGroup show that perception has the lowest average accuracy among the three tasks. Correct perception answers are generally associated with higher reasoning and summarization accuracy within the same QA groups. We release the dataset and annotation framework to support reproducible evaluation of video understanding across tasks within shared object contexts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.