acceptodds
Under review as a conference paper at ICLR 2027

ReelBench: Towards Systematic Evaluation of Multi-Shot Narrative Video Generation under Image Reference

Abstract

Multi-shot narrative video generation increasingly relies on image references, including asset and storyboard images, to maintain visual and narrative consistency across independently generated shots. Yet existing benchmarks primarily assess final video quality or rank model capabilities, offering limited insight into how different forms of image conditioning affect generation. We introduce ReelBench, an evaluation framework and dataset for image-guided multi-shot narrative video generation. ReelBench comprises 300 story-style pairs with associated asset and storyboard images across three visual styles: Pixar-style animation, live-action realism, and 2D anime. It supports controlled comparisons among image reference conditions within the same story and evaluates generated videos along seven dimensions and 40 subdimensions: instruction alignment, character quality, visual quality, physical plausibility, visual continuity, cinematic quality, and narrative coherence. Our experiments show that stronger visual anchoring through first-frame conditioning does not consistently improve overall quality, and that reference complexity affects physical plausibility more strongly than overall scores. ReelBench agrees with human overall preferences in over 90% of pairwise comparisons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.