Getting It Right, Under Constraint: Evaluating and Improving LLM Creativity
Abstract
One psychological account defines creativity as *constrained problem solving*, shaped by five situational factors: pressure, pleasure, opportunity, prototypes (earlier works to imitate and transform), and helpers. Current creativity benchmarks for large language models (LLMs) vary none of these factors. Their evaluation either scores correctness with the conventional route open or closes it without comparing against an unconstrained counterpart, and we find it is unstable: swapping incidental numbers alone costs models 8.9 to 15.2 points on OMEGA. We therefore build OMEGA-Constraint and Omni-MATH-Constraint, pairing each problem with versions that ban the reference method, cap length, or forbid a character, closing the conventional route while another stays open. We then add tool access as opportunity, threats as pressure, and rewards as pleasure. Whereas constraint can sharpen human creativity, LLMs lose accuracy as constraints accumulate (under all three, gpt-5.5 drops from 97.5 to 49.0 percent on OMEGA) and reapply banned methods under new names. As in humans, tools help every model, but threats and rewards of up to a million dollars move none by more than 6 points. Our proposer-composer harness, based on stage theories of creativity, supplies the remaining two factors: a small proposer (the helper) brainstorms candidates (the prototypes), which a larger composer develops into its own answer. With the best of ten proposers, both 32B composers gain in every condition, and a single 3.8B proposer lifts gpt-4.1 from 9.0 to 25.0 percent under all three constraints on OMEGA. On constrained math, our harness with a String Seed of Thought-prompted proposer beats plain and fine-tuned proposers in 17 of 28 conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.