acceptodds
Under review as a conference paper at ICLR 2027

Autonomous Weighted Schema Direct Preference Optimization for Materials Discovery

Abstract

Generative artificial intelligence can create candidate molecules and materials for tasks in the physical and biological sciences, but it can be limited in producing better candidates along characterized properties than those of the systems in the training data. Reinforcement learning, in particular Direct Preference Optimization (DPO) on chemical language models (CLMs), can shift output chemical distributions in desirable ways, but constructing preference datasets in an autonomous manner presents further challenges. Here, we outline a novel methodology for multi-objective optimization where a defined target region of chemical properties is compared to generated distributions and used for the updating of preference weights that partition generated molecules into a cohesive, multi-object preference dataset. In this process, termed autonomous weighted schema (AWS) DPO, we observe enhancements to generated distributions beyond that observed without autonomous weight updating as well as reward functions that put the multi-objective frontier into a single score, and we thereby present a general algorithm for tuning generative models for materials discovery.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.