TOC-Bench: Diagnosing Temporal Object Consistency in Video Large Language Models
Abstract
Video large language models (Video-LLMs) have achieved remarkable progress in general video understanding, yet their ability to maintain temporal object consistency remains insufficiently evaluated. This capability connects the identity, state, and persistence of the same object across occlusion, disappearance, reappearance, and cross-object interactions. We introduce TOC-Bench, a diagnostic benchmark that grounds temporal questions in explicit object tracks and structured event timelines. Candidate answers are fixed by object-event records before language realization. A three-layer temporal-necessity filtering protocol probes text-only, single-frame, and frame-shuffled shortcuts, removing 60.7% of 45,527 candidates. From the resulting pool, expert verification yields 2,323 QA pairs over 1,951 real videos, spanning 10 diagnostic dimensions and four deterministic task formats. Experiments on 26 Video-LLMs show substantial difficulties in repeated-event counting, exact event ordering, and verification of unsupported object/event premises: the strongest evaluated model achieves 47.1% accuracy, versus a three-person human average of 89.3%. Controlled 16/32/64-frame comparisons show that a larger uniform frame budget alone does not consistently improve performance. TOC-Bench offers new instance-linked annotations and a focused platform for diagnosing and improving object-aware temporal reasoning. The resource is available at https://anonymous.4open.science/r/toc_bench-1355.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.