Measuring Combinatorial Creativity in LLMs
Abstract
Combinatorial creativity, the process of combining familiar ideas in unfamiliar ways, is one of the foremost engines of scientific, technological, and artistic progress. While large language models (LLMs) can fluently produce novel text relative to their training data, their capacity to combine information in novel *and* useful ways, in order to invent new concepts, remains unclear. This work introduces **Kombine**, a combinatorial creativity benchmark tasking models to discover surprising, original, and useful associations, analogies, and blends between unrelated entities, and to use those analogies and blends to *invent novel concepts by re-combining the parts of existing ones*. In a large-scale study across 35 LLMs, we find that: **(1)** LLMs are 10-15 percentage points better at discovering associations and analogies than blends, and **(2)** exhibit an *abstraction bottleneck* on blending, in which even frontier LLMs struggle to find a shared abstraction between inputs and instantiate a coherent blended space. **(3)** LLMs also exhibit "inventive homogeneity," in which one in five concepts invented for a given item reuses properties invented by another model. Moreover, **(4)** models are more likely to reuse other invented properties on blends than analogies, and **(5)** models from the same provider are more likely to reuse properties than those from different providers. Our results reveal creative strengths in LLMs' capacity for association and analogy, point to an abstraction bottleneck on blending, and show how output homogeneity constrains the conceptual diversity of inventions. In summary, we provide a tool that can be used to assess progress towards more creative LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.