DySCo: A Lightweight Data Curation Framework for Reinforcement Learning in SMILES-based Molecular Design
Abstract
Molecular design aims to generate chemicals with desired properties, but the vastness of chemical space makes exhaustive search impractical. Reinforcement learning (RL)-based generative models have shown promise by enabling self-guided exploration. However, standard RL fine-tuning faces an trade-off: aggressive reward maximization often leads to mode collapse with loss of structural diversity, while maintaining diversity typically comes at the cost of slow convergence and high computational overhead due to extensive sampling. To address this challenge, we introduce DySCo (Dynamic State-adaptive Coreset), a lightweight and task-agnostic framework for adaptive data curation during RL training. Unlike static replay buffers, DySCo treats the pre-trained prior policy as a continual source of structurally diverse molecules. It dynamically monitors the contribution of the current policy to high-reward elites as a simultaneous signal of diversity loss, and adaptively constructs compact corrective coresets that balance structural novelty and reward optimization. Empirical results on challenging multi-objective molecular generation tasks demonstrate that DySCo reduces per-epoch computational cost by 20%-45% while achieving competitive or superior reward scores and significantly higher structural diversity and improved sample efficiency compared to existing baselines. As a data-level method, DySCo is complementary to reward-level and policy-level diversity approaches and compatible with different RL optimizers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.