VAHBench:A Harness Benchmark for Agentic Long-Video Generation
Abstract
AI agents are commonly used for everyday information retrieval and coding. Weask whether they can also interpret human requests and turn them into complete videos. We introduce VAHBench, a benchmark that tests this ability through fiveshort-drama briefs specifying story, characters, dialogue and duration. Agentsplan and carry out production using shared generation tools. Evaluation combines video completeness and basic quality with checks of production reliability and reporting accuracy. Across 240 runs covering 48 configurations of general-purpose and video-specific agents, higher quality scores do not necessarily mean more complete delivery. Missing episodes, repeated footage and static filler reveal gaps between producing a video file and fulfilling the brief. Added composition tools help some agents but do not consistently improve performance. VAHBench mea-sures how well agents translate user intent into finished video content and helpsidentify where that process breaks down.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.