ScreenWritingBench: Evaluating LLMs Across the Screenwriting Workflow with Dimension-Specific Evaluation
Abstract
Professional screenwriting develops through planning, drafting, and revision across several artifact types. Existing writing benchmarks sample only parts of this workflow, making it difficult to assess whether large language models (LLMs) can support the successive tasks of screenwriting. Dramatic quality also lacks an objective ground truth and calls for a different form of judgment from task completion or artifact compliance. To address these gaps, we introduce ScreenWritingBench, a collection of 279 tasks constructed and validated by professional screenwriters. The tasks span artifacts from proposals to scripts and operations from initial generation to revision; 219 include dialogue history. We evaluate responses along five dimensions: Task Completion, Contextual Consistency, Artifact Compliance, Basic Lexical Defects, and Dramatic Quality. The first four use rubrics, whereas Dramatic Quality places each response relative to eight expert-ordered anchors. An evaluation of 18 models reveals substantial variation across dimensions. We report both an overall leaderboard and dimension-level scores to characterize these differences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.