acceptodds
Under review as a conference paper at ICLR 2027

ReelOnce: Adaptive Reference Policy and Evaluation-Driven Hierarchical Refinement for Agentic Multi-Shot Story Video Generation

Abstract

Multi-shot story video generation requires shot-specific visual guidance while preserving consistency across characters, scenes, and narrative events, often requiring repeated manual revision. Existing workflows struggle to determine whether a problem originates in story planning, reference construction, or video generation. Fixed-stage retries are costly and may miss the root cause or disrupt content that is already correct. We present ReelOnce, an agentic framework that combines an adaptive reference policy with evaluation-driven hierarchical refinement. It selects an appropriate reference format for each shot, distinguishes the state depicted in a reference from the transition to be generated, and uses evaluation findings to identify the smallest necessary repair scope. Persistent story memory preserves narrative constraints and unaffected decisions throughout refinement. We further introduce Reconstructed Story Semantic Alignment (RSSA), which measures semantic consistency with the source story by reconstructing observable events from generated videos. On Pixar-style videos generated from 20 diverse stories, ReelOnce achieves strong RSSA performance while remaining competitive across VBench quality dimensions, demonstrating that the framework can balance story fidelity with visual quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.