SCORED: Selective OSDA Generation with COntrastive REwards from Denoising
Abstract
Scientific design often requires candidates that favor a target condition over plausible alternatives. Conditional generative models do not directly optimize this preference, while evaluating every candidate across multiple conditions is often too costly for online optimization. We introduce SCORED, a model-intrinsic selectivity reward that uses shared noise to contrast a conditional diffusion generator's denoising errors for the same candidate under target and competitor conditions. Scoring needs only forward passes: a six-condition panel takes 7.86 ms, against 8.3 s for a docking-based evaluator. We apply SCORED to generate organic structure-directing agents (OSDAs), molecules that guide zeolite framework formation. Mean-centered reward-weighted regression incorporates offline binding information, after which multi-objective reinforcement learning (RL) combines selectivity and molecular-quality rewards without physical evaluation in the online reward loop. Across three runs, SCORED reaches a mean competition score of kJ/mol Si over six held-out frameworks under frozen-pose force-field evaluation, where a recent LLM-based generator scores 9.72 (lower is better). Matched ablations that drop the competitor contrast (6.34) or the denoising reward entirely (6.50) are both worse, isolating the contribution of the target–competitor comparison. Extensions to structure-based drug design and metal–organic framework generation give uneven gains across models and domains, which we examine through pre-RL score–property alignment. These results support conditional denoising comparisons as inexpensive online feedback for selectivity optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.