acceptodds
Under review as a conference paper at ICLR 2027

ScreenWritingBench: Evaluating LLMs Across the Screenwriting Workflow with Dimension-Specific Evaluation

Abstract

Professional screenwriting develops through planning, drafting, and revision across several artifact types. Existing writing benchmarks sample only parts of this workflow, making it difficult to assess whether large language models (LLMs) can support the successive tasks of screenwriting. Dramatic quality also lacks an objective ground truth and calls for a different form of judgment from task completion or artifact compliance. To address these gaps, we introduce ScreenWritingBench, a collection of 279 tasks constructed and validated by professional screenwriters. The tasks span artifacts from proposals to scripts and operations from initial generation to revision; 219 include dialogue history. We evaluate responses along five dimensions: Task Completion, Contextual Consistency, Artifact Compliance, Basic Lexical Defects, and Dramatic Quality. The first four use rubrics, whereas Dramatic Quality places each response relative to eight expert-ordered anchors. An evaluation of 18 models reveals substantial variation across dimensions. We report both an overall leaderboard and dimension-level scores to characterize these differences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.