acceptodds
Under review as a conference paper at ICLR 2027

CreativitySuite: A Comprehensive Framework for Evaluating and Enhancing Creativity in Large Language Models

Abstract

Creativity is widely regarded as a general capability of intelligence, yet its evaluation in large language models (LLMs) remains fragmented across domains and often conflates effectiveness with originality. We operationalize the standard definition of creativity through response-level quality and novelty, together with task-level creative space. Controlled experiments demonstrate that separate evaluation of quality and novelty improves the consistency and reference agreement of quality judgments, and that expanding creative space substantially increases the proportion of high-quality, high-novelty responses. Building on this formulation, we introduce CreativitySuite, a creativity training and evaluation framework spanning seven domains, with 19,858 training samples, 21,220 evaluation samples, and six complementary evaluation methods at both domain and instance levels. Experiments show that in-domain supervised fine-tuning yields localized gains, whereas joint training across domains yields inconsistent transfer and tends to improve correctness at the expense of novelty. Moreover, the measured effects of training vary considerably across evaluators. These findings indicate that imitating high-quality responses is insufficient for general creativity, which requires explicit coordination of quality and novelty. Once permitted, we will release our data and code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.