ReVi: Recursive Recap Video Generation
Abstract
Novels tell rich stories, but their textual form cannot directly present the sights and sounds of the worlds they describe. Film adaptations offer an audiovisual experience but require substantial production effort and viewing time; film recaps are shorter but depend on existing footage. Neither route directly provides a concise audiovisual account of a novel. We therefore study novel-to-recap video generation: producing a self-contained, source-grounded audiovisual retelling directly from a novel within a fixed viewing-time budget. This task requires compressing the story while preserving causal connections, allocating sufficient time to retained events, and aligning narration with newly generated visuals. We propose ReVi, a dual-loop multi-agent framework for Recursive Recap Video Generation. Its Inner Loop plans story content, narration time, and corresponding shots before video generation. Across novels, its Outer Loop performs recursive self-improvement by revising both the editing policy and the meta policy that governs its revision. Experiments on two evaluation sets show that ReVi improves recap quality and duration compliance over representative baselines, with human ratings and multi-round improvement results providing further support. Our code is available at https://anonymous.4open.science/r/ReVi-84EF.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.