acceptodds
Under review as a conference paper at ICLR 2027

DramaChain Bench: An End-to-End Benchmark for Short-Drama Generation

Abstract

Commercial short-drama production spans multiple stages. Most benchmarks, however, assess video generation in isolation, leaving unclear whether defects originate in the evaluated stage or propagate from earlier ones. In addition, most use manually constructed inputs that do not reflect real production conditions. We present DramaChain Bench, the first short-drama benchmark evaluating pipeline-native outputs from storyboard design to multi-episode delivery. It combines an evaluation schema with three systems for generation, human annotation and automated judging. DramaChain Dimensions defines 63 leaf dimensions across six granularities and five axes: input fidelity, internal consistency, generation plausibility, visual quality and cinematic expressiveness. From supplied scripts, DramaChain Agent generates storyboards, keyframes, shot videos and finished dramas, with its workflow and output quality benchmarked against commercial platforms. DramaChain Labeling System provides three independent assessments for each of 5,569 reference items, with spatial and temporal defect localisation where applicable. DramaChain Agentic Judge combines a master–sub-agent architecture, 29 specialist modules and iterative tool use, achieving a mean model-level PLCC of 0.918 against human scores. Using this framework, we evaluate 37 generation models: 20 text LLMs, 8 image models and 9 video models. A series of judge models trained on these annotations with supervised fine-tuning and GRPO further achieve higher agreement with human ratings than the prompted agentic judge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.