acceptodds
Under review as a conference paper at ICLR 2027

Autonomous Discovery of Training Distribution for Controllable Molecular Generation

Abstract

Molecular generative models are typically pretrained on fixed drug-like datasets, implicitly constraining the structural support of the learned representation and limiting downstream adaptability. We investigate whether pretraining distribution support acts as a controllable lever on generative expressivity, and whether expanding this support through structurally distant molecules improves downstream diversity–activity trade-offs. We instantiate this hypothesis using a large language model (LLM) as an autonomous research agent and a lightweight autoregressive Transformer as the molecular generator. The agent iteratively proposes molecular pretraining mixtures, the generator is retrained under fixed model capacity and compute budget, and performance is evaluated on the task of spleen tyrosine kinase (SYK) inhibitor generation and optimization using a joint objective combining scaffold entropy and predicted pIC after fine-tuning on known SYK inhibitors. Across multiple iterations, the search identifies boundary configurations composed of molecules maximally distant from canonical drug-like datasets, yielding higher diversity–activity scores than conventional baselines. The best-performing configuration improves the joint diversity–activity objective by 1.6% relative to the baseline, driven primarily by increased scaffold entropy while maintaining competitive predicted affinity. These results suggest that training distribution discovery can serve as a data-centric mechanism for improving controllable molecular generation under constrained model capacity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.