acceptodds
Under review as a conference paper at ICLR 2027

NEXUS: A Causal-Aware Framework for Multimodal Long-Form Movie Understanding

Abstract

Understanding long-form videos, especially movies, requires jointly integrating visual, audio, and dialogue signals while persistently tracking character dynamics and event causality. Agentic models are increasingly used to tackle this challenge. However, these approaches either rely on subtitles as a stand-in for video, leaving other modalities underutilized, or reason directly over raw frames at a computational cost that scales sharply with video length and query complexity. Moreover, the field is largely evaluated through multiple-choice (MC) benchmarks, which do not always demand the depth of reasoning that free-form responses require. Under this lens, we introduce NEXUS, a framework for long-form video understanding built around three core components. First, we propose NEXUS-MF, an end-to-end multimodal understanding pipeline that processes long videos into a structured context representation capturing character dynamics and long-term causal dependencies. Second, we present NEXUS-OEBench, a large-scale, human-annotated, open-ended QA benchmark comprising 1,701 question–answer pairs spanning 150+ hours of video content, with questions demanding deep causal and multimodal reasoning. Third, we introduce NEXUS-QA, a video QA agent that causally reasons over NEXUS-MF's structured representation, actively seeking and verifying evidence to construct complete causal chains before forming a final response. NEXUS outperforms baselines on QA benchmarks with 3.4× fewer tokens and 3.3× lower costs, while maintaining stable computational overhead as video length increases. Additionally, NEXUS-MF's multimodal representation generalizes well on the context-retrieval task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.