acceptodds
Under review as a conference paper at ICLR 2027

Constructive Divergence: A Judge-Free Reward for Creative Language Generation

Abstract

Creativity is a core capability for autonomous research agents, yet optimizing it remains challenging: verifiable reward signals are scarce, human annotations and LLM-as-a-judge can be costly or contaminated, and heuristic rewards are vulnerable to hacking. To address this challenge, we propose Constructive Divergence, a scalable creativity score requiring only the token probabilities of two reference models. Grounded in the psychological view that creativity requires both novelty and value, we operationalize novelty as the surprisal of an output and value as the recoverability of its original intent. Concretely, we compute surprisal as the negative log-likelihood of the output under a strong reference model, and recoverability as the conditional log-likelihood of the prompt given the output under a second reference model. Notably, Constructive Divergence serves as both a reliable evaluation metric and a cost-effective training reward. As an evaluator, Constructive Divergence is deterministic and robust to reference-model choice, with 70% agreement across model configurations. On paper-abstract generation, optimizing models with DPO and GRPO using Constructive Divergence improves creativity with an average increase of 25%. Our trained models achieve creativity gains up to 4.0 times larger than the average gain of the baselines, and their outputs are preferred by LLM-as-a-judge models in up to 61% of pairwise comparisons, as further examined with human annotators.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.