acceptodds
Under review as a conference paper at ICLR 2027

Verify What Matters: Cost-Aware Long Video Generation with GIVER

Abstract

Although recent agentic systems can generate long multi-shot videos from storyboards, physical inconsistencies can persist, such as fingers passing through a cup during lifting. Repairing these defects requires identifying failures and guiding regeneration without altering the requested content. To this end, we introduce GIVER, an agentic framework that embeds verification-and-repair loops into the generation process. Specifically, GIVER constructs a fixed Subject–Object Interaction Graph (SOIG), whose nodes encode storyboard-derived visual requirements and whose edges encode their logical prerequisites. Using this graph, GIVER verifies reusable assets and repairs them when needed, then runs two local loops: L1 checks the first frame’s initial state, while L2 checks required video behavior and visible artifacts, retaining the selected first frame. Failed-node evidence guides targeted revisions to generation fields, and each new candidate is rechecked against the same requirements. On ViMax-Bench and ST-Bench-10, GIVER leads the compared pipelines in fidelity and physical plausibility under automated and blinded human evaluation, with automated severe-artifact rates of 0.70% and 2.80%, respectively. Additionally, our ablation study shows improved quality with approximately 24–25% fewer video generation and editing operations than two flat-verification configurations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.