acceptodds
Under review as a conference paper at ICLR 2027

Shape the Narrative: A Dataset for Evaluating Message-Driven Narrative Video Editing

Abstract

Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative derived from messages the editor wishes to convey. Benchmarks for a closely related task, video summarization, reduce editorial intent to a single, message-agnostic notion of saliency and thus do not account for this diversity. For evaluating message-driven video editing, we present MEDit-Bench, a dataset and benchmark, which pairs long-form videos with multiple editing messages and multiple professionally produced edits per message, demonstrating that different messages yield substantially different edits from the same source. We define an automatic evaluation protocol based on temporal alignment metrics, and find that an LLM-as-a-judge preference, a natural proxy for narrative quality, is unreliable for this task due to severe position bias. We additionally annotate each message with ambiguity and contextfulness scores, and show that both dimensions negatively correlate with model performance, establishing video editing difficulty as a meaningful stratification factor. Experiments with state-of-the-art MLLMs and reinforcement fine-tuned baselines show that while strong models approach human temporal alignment at lenient thresholds, all models fall behind humans at stricter criteria. A human perceptual study further shows that professional human edits remain preferred over model outputs on every evaluated aspect.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.