acceptodds
Under review as a conference paper at ICLR 2027

Beyond Isolated Clips: Benchmarking Inter-Segment Narrative Continuity in Movie-Level Video Understanding

Abstract

Recent advancements in multimodal large language models (MLLMs) have extended video understanding from short clips to long-form cinematic content. However, existing benchmarks predominantly evaluate isolated video segments, overlooking the narrative continuity essential to full-length movies, such as cross-scene information transfer, character consistency, and plot foreshadowing. To bridge this gap, we introduce CineScriptBench, the first movie-level benchmark explicitly designed to evaluate inter-segment historical consistency in cinematic script generation. Comprising 73 full-length movies (153 hours), 5,524 scenes, and 56,993 fine-grained evaluation rubrics, CineScriptBench is constructed via a novel five-stage automated pipeline that injects rolling historical context into scene-level annotations. We propose a comprehensive evaluation protocol measuring both temporal localization and content accuracy, featuring a dedicated dimension for historical consistency via an LLM-as-a-Judge framework. Extensive evaluations of frontier MLLMs reveal severe limitations in long-form narrative understanding. Our findings highlight the critical gap between current MLLM capabilities and true movie-level comprehension, establishing a rigorous baseline for future research in cinematic narrative reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.