acceptodds
Under review as a conference paper at ICLR 2027

Video Text Understanding Across Diverse Scenes: An Evidence-Guided Benchmark

Abstract

Video text understanding involves interpreting on-screen text together with its visual context. Audits of existing benchmarks have identified questions answerable from a single frame or by text extraction alone, motivating broader assessment of this joint understanding. We introduce an evidence-guided benchmark for video text understanding across diverse scenes. Its 16,165 evaluation items comprise multiple-choice questions and true/false judgments over 2,629 videos, organized into 10 primary scene categories and 34 fine-grained subcategories. The construction workflow plans structured evidence before generating questions, followed by automated verification, text-only filtering, and human correction. We evaluate 26 open-weight multimodal checkpoints spanning 1B–235B parameters. On the matched multiple-choice roster, mean accuracy is 55.15% with video versus 41.49% with questions and options alone. Paired judgments reveal difficulty both accepting supported statements and rejecting their counterfactual counterparts, while performance varies across scenes. We further manually review a subset of 500 base questions and their associated variants for controlled evidence diagnostics: at equal frame counts, sampling distinct evidence windows outperforms repeating one window center on average. The benchmark provides a testbed for evaluating joint text–visual understanding and diagnosing remaining model weaknesses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.