acceptodds
Under review as a conference paper at ICLR 2027

VAHBench:A Harness Benchmark for Agentic Long-Video Generation

Abstract

AI agents are commonly used for everyday information retrieval and coding. Weask whether they can also interpret human requests and turn them into complete videos. We introduce VAHBench, a benchmark that tests this ability through fiveshort-drama briefs specifying story, characters, dialogue and duration. Agentsplan and carry out production using shared generation tools. Evaluation combines video completeness and basic quality with checks of production reliability and reporting accuracy. Across 240 runs covering 48 configurations of general-purpose and video-specific agents, higher quality scores do not necessarily mean more complete delivery. Missing episodes, repeated footage and static filler reveal gaps between producing a video file and fulfilling the brief. Added composition tools help some agents but do not consistently improve performance. VAHBench mea-sures how well agents translate user intent into finished video content and helpsidentify where that process breaks down.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.