Generating Materials with Excellent Properties via Actor-Critic Diffusion
Abstract
Diffusion models, as the basis for state-of-the-art foundation models for materials discovery, potentially signal a paradigm shift from the expensive, manual trial-and-error industrial practice. Technology across multiple fields is powered by materials with desirable and often rarely seen properties. Classifier-free (or guided) diffusion has been the main approach in generating materials with these conditioned properties, however, these approaches tend to fall back to existing materials when conditioned on extreme property values. This poses significant hurdles in discovering novel, stable materials with excellent properties. Instead of formulating the problem as conditional diffusion, we define a novel Reinforcement Learning (RL) process for the reverse diffusion to maximize property values as rewards. RL explores the composition space freely, hence not bounded by existing materials in the training dataset. We present CrystRL, a reinforcement learning-based method that can generate novel, stable materials with excellent properties. CrystRL fine-tunes the unconditional reverse diffusion using an actor network, which suggests denoising moves to good materials. It is supplemented by a critic network, which predicts the best property values from a noisy state. Experimental results show that CrystRL discovers novel and stable crystals with out-of-training-distribution ML-predicted bulk moduli of up to 408 GPa and DFT-verified bulk moduli of up to 397 GPa more consistently than classifier-free guided generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.