acceptodds
Under review as a conference paper at ICLR 2027

DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluation

Abstract

Long-form video generation is rapidly moving from short, single-scene synthesis toward minute-long, multi-shot creation with narrative structure, cinematic control, audio, and cross-modal synchronization. However, evaluating such videos remains challenging, since existing benchmarks largely focus on local visual quality, short-horizon temporal consistency, or generic prompt alignment, and provide limited diagnosis of workflow failures and user-dependent preferences. We introduce DirectorBench, a personalized multi-agent diagnostic benchmark for long-form video generation. DirectorBench evaluates generated videos with respect to 80 structured metadata entries, 7 user profiles, and 40 checkpoint criteria across 5 dimensions: script, visual, audio, cross-modal, and stability. Instead of reducing quality to a single aggregate score, DirectorBench localizes checkpoint-level bottlenecks and supports profile-aware evaluation. We evaluate 4 long-form video generation workflows, 6 base LLMs, and 7 user profiles. Across workflows, DirectorBench reveals a between-unit bottleneck: transition quality averages only 0.256 and reaches 0.356 for the best workflow, while prompt-level user demand fulfillment averages 0.71. A study with 14 annotators yields case-level Spearman correlation 0.739 and 84.3% pairwise ranking accuracy; the human-rated bottleneck appears in the automatic top-three list in 96.5% of cases. Additional experiments show that profile-conditioned prompts change speech density and video characteristics. These findings highlight the importance of diagnostic and profile-aware benchmarking for agentic long-form video generation. Our code is available at https://anonymous.4open.science/r/DirectorBench-835E/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.