acceptodds
Under review as a conference paper at ICLR 2027

TwistBench: Benchmarking Transformational Creativity in LLMs via Literary Plot Twists

Abstract

Benchmarks and tests of creativity for Large Language Models (LLMs) are now abundant—it has been shown that LLMs can realize aspects of combinatorial creativity, exploratory creativity, and that they exceed human performance on psychometric tests of creativity. However, these evaluations have failed to address more complex types of creativity recognized in cognitive science, such as *transformational creativity*. Widely held to underpin major scientific, artistic, and technological breakthroughs, transformational creativity describes the ability to *radically modify an existing conceptual space*, such as a scientific paradigm, a musical genre, or, as we show, the plot of a short story. To resolve this gap in the benchmarking landscape, we introduce **TwistBench**, a benchmark of literary plot twists for assessing transformational creativity. We collect a set of stories with plot twists written by famous human authors and compare them against stories written by 71 LLMs, across scoring dimensions *diversity, surprise, coherence*, and *realism*. Our main findings are that: (1) Even the latest frontier LLMs underperform expert humans in transformational creativity, though elaborate prompting strategies like in-context regeneration can raise the performance of select models to match expert humans. (2) Two dominant failure modes plague LLMs when applied to open-ended, transformative tasks: (a) *mode collapse*, where frontier models are capable of generating high quality twists, but these twists are largely the same and lack diversity; and (b) *breaking the world model,* where models give surprising and coherent plot twists, but do so in a way that measurably departs from physical reality (e.g., time travel, ghosts), violating literary "fair play." (3) An analysis of the creative processes employed by reasoning models reveals process-level homogeneity among LLMs. Namely, models always choose the plot twist *first*, then retroactively determine the plot conditioned on the twist. In summary, while today's LLMs struggle to surpass the transformational creativity of expert humans, we introduce a task and framework to assess progress towards this ability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.