Self-Reinforcing Creativity in Diffusion Generative Models
Abstract
Creativity is commonly understood as the production of content that is both novel and meaningful. While existing diffusion models can synthesize images that appear to generalize beyond the training data, their creative capacity remains limited. In this paper, we present a method for fine-tuning a diffusion model to generate highly creative image samples that extend beyond its initial output distribution. This is achieved by rewarding the model for generating images that are perceptually distant from its original outputs, using a distance function derived from the model itself. At the same time, we keep its distribution close to the pre-trained one through a KL penalty. Despite using no external supervision or reward functions, our model preserves the pre-trained model's image quality while producing highly novel samples. We further show that existing reinforcement-learning methods for diffusion-model post-training can fine-tune such models directly, as they optimize KL-regularized reward objectives of the same form. These off-the-shelf methods can therefore use the model's own novelty signal to amplify its creative capacity, making the process self-reinforcing. We demonstrate our method across settings ranging from a simple two-dimensional distribution and unconditional MNIST generation to a large-scale text-to-image model. Across these domains, the fine-tuned models produce outputs that depart from the typical generations of their pretrained counterparts while preserving coherent structure. The resulting models generate such outputs through standard sampling, eliminating the need for per-sample optimization or specialized inference-time steering.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.