Learning Deployable Policy Improvements from Measured Molecular Panels
Abstract
A molecular policy can earn a higher reward yet retain little of that improvement once endpoint regression is controlled. We study this gap in finite ADMET panels: historical measurements cover every candidate, while deployment exposes only molecular structures and endpoint directions. DeltaADMET-R1 combines an exact, anchor-relative multi-endpoint target with a total-variation (TV) controller that bounds regression on each endpoint. Its direct learning algorithm, Contextwise Primal–Dual Policy Optimization with structured Reasoning (CPDPO-R), differentiates the expected utility of the controlled policy using complete action feedback. When the TV budget is active, this objective maximizes gain per unit of policy displacement. We establish target feasibility, contraction under identityaligned permutation averaging, and a tight outcome-free endpoint bound; source calibration provides a less conservative statistical operating point. On scaffolddisjoint RLM clearance and solubility panels, five-seed Qwen3-8B CPDPO-R exceeds matched scalar projection by .0000731 (95% CI [.0000065, .0001403]) under the same ϵ = .005 budget. A post hoc joint source-conformal analysis retains mean gain .000769, versus .000641 for scalar, at nominal 95% coverage. The target-based pipeline yields positive aggregate gains across six Biogen, TDC, and OpenADMET evaluations, with further analyses of larger panels, four endpoints, and missing measurements. Qwen4B and Mistral show that higher raw reward need not translate into a controlled advantage. These results support retained utility as a useful RL objective, with model-family transfer as the principal empirical boundary.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.