AVID: Action-Value Guidance and Identity-Aware Learning for Candidate-Set Molecular Discovery
Abstract
Molecular optimization rewards individual molecules, but a finite-budget discovery campaign seeks a growing set of distinct, high-quality candidates. Repeatedly generating a high-scoring molecule can maintain high per-output reward without adding a candidate. We introduce AVID, which couples action-value guidance during generation with identity-aware learning after completion. A frozen, archive-trained prefix-action guide biases token selection toward high-reward continuations. Once a molecule is complete, an identity ledger assigns within-run repeats zero weight in the direct policy loss while preserving cached rewards and the optimizer's other learning terms. Across all 23 tasks of the Practical Molecular Optimization (PMO) benchmark and five seeds per task, AVID acquires an average of 7,114.1 distinct candidates meeting fixed archive-derived reward thresholds within 32,768 generation attempts and at most 10,000 new oracle calls. This is 45.5% more than Augmented Memory, with higher task means on 18 of 23 tasks. Component studies distinguish the two interventions: adding guidance to identity weighting improves mean and top-10 reward at 512 newly scored identities on every task, while long-horizon experiments expose the trade-off between candidate quality and acquisition. Guidance can raise the qualifying fraction yet reduce distinct yield if new identities arrive more slowly. These results show that reward quality and qualifying-set growth provide complementary views of molecular optimization under finite budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.