Active Molecular Query Synthesis by Influence-guided Latent Search
Abstract
Molecular screening relies on accurate property predictions, but obtaining the data to train these predictors can require costly assays or computations. Active learning reduces this cost by selecting molecules for evaluation, but its choices are often restricted to an existing library. Molecular generation allows the learner to construct additional queries and choose them for their estimated effect on prediction. We propose Influence-guided Molecular Query Synthesis (iMQS) to computationally generate molecular structures for their value as training data under a limited evaluation budget. Its acquisition objective is the estimated reduction in prediction loss on a validation set. An influence function approximates this improvement without retraining for every candidate and guides search in a continuous molecular representation. Our analysis bounds errors in comparing candidate improvements and the validation improvement lost through query selection when the predictor's final layer is refitted. Across nine molecular endpoints, iMQS improves predictive learning over several pool and generation baselines. In two AutoDock Vina applications, mean recovery of measured screening actives increases over the acquisition and refitting procedure. These results support molecular query generation as a way to acquire useful training data under limited evaluation budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.