TriCreativity: Benchmarking and Eliciting Combinational, Exploratory, and Transformational Creativity for Large Language Models
Abstract
Creativity is an important capability for large language models in scientific discovery, technological innovation, and open-ended problem solving. Existing benchmarks span multiple domains, but leave gaps in distinguishing the creative demands of tasks and assessing departures from existing conceptual frameworks. To this end, this paper introduces TriCreativity, a benchmark combining Boden’s combinational, exploratory, and transformational taxonomy with multidimensional evaluation. We also introduce Counterfactual Transformation (CFT), a dataset that supplies counterfactual changes to fundamental premises. Models must develop additional representations and organizing principles to construct coherent and useful conceptual frameworks under the changed premises. Across 11 large language models spanning multiple families and sizes, creative strengths vary by creative demand and domain: even within exploratory creativity, mathematics, coding, and writing favor different models. Within-family scaling yields uneven gains. Most evaluated generic reasoning procedures improve selected tasks or metrics but struggle to achieve joint gains across creative objectives. Thus, we further introduce Taxonomy-Guided Creativity Elicitation (TCE), a training-free method that coordinates generation, search, selection, and verification around task-specific creative requirements and achieves the best aggregate performance among the compared procedures. Novelty and appropriateness are strongly associated in CFT but weakly related in constrained writing. Together with output-level cases, this suggests that when familiar premises fail, effective problem solving may depend not only on finding alternative answers but on constructing new representations and organizing principles.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.