Autonomous Weighted Schema Direct Preference Optimization for Materials Discovery
Abstract
Generative artificial intelligence can create candidate molecules and materials for tasks in the physical and biological sciences, but it can be limited in producing better candidates along characterized properties than those of the systems in the training data. Reinforcement learning, in particular Direct Preference Optimization (DPO) on chemical language models (CLMs), can shift output chemical distributions in desirable ways, but constructing preference datasets in an autonomous manner presents further challenges. Here, we outline a novel methodology for multi-objective optimization where a defined target region of chemical properties is compared to generated distributions and used for the updating of preference weights that partition generated molecules into a cohesive, multi-object preference dataset. In this process, termed autonomous weighted schema (AWS) DPO, we observe enhancements to generated distributions beyond that observed without autonomous weight updating as well as reward functions that put the multi-objective frontier into a single score, and we thereby present a general algorithm for tuning generative models for materials discovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.