Long-Horizon Movie Reasoning: A Scalable Benchmark and Structured Memory Framework
Abstract
Long-form video understanding requires models to reason consistently over extended temporal contexts, yet existing benchmarks provide limited evaluation of long-range memory and high-level narrative reasoning. We introduce a semi-automated pipeline for constructing long-form video reasoning benchmarks from full-length movies and their recap narrations, reducing annotation cost while preserving reliable temporal and semantic alignment. Based on this pipeline, we construct Movie Reason Benchmark (MRBench), which evaluates holistic reasoning over long videos with narrative semantics at multiple levels of abstraction. We further introduce complementary metrics for temporal localization under multiple tolerance thresholds and semantic alignment across narrative granularities. To address the challenges exposed by MRBench, we propose EdgeMemory, an edge-centric graph-structured memory framework that decomposes complex reasoning tasks, organizes multi-level semantic representations, and progressively refines their relations through a dedicated reasoning pipeline. Extensive experiments show that EdgeMemory consistently improves long-form video reasoning across diverse models. Further analysis reveals distinct capability boundaries and failure modes across model scales, highlighting persistent challenges in long-range evidence integration and abstract narrative reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.