acceptodds
Under review as a conference paper at ICLR 2027

Learning to Evolve Molecules with Local-Global Credit Assignment

Abstract

Large language models (LLMs) optimize molecules through trajectories of successive proposals and evaluator feedback, seeking property improvements under chemical validity and similarity constraints. Outcome-level reinforcement learning trains on these trajectories but assigns the same advantage to every turn, despite differences in molecular quality. We introduce Learning to Evolve (L2E), which trains LLMs through a local-global principle. For trajectories sharing an initial molecule and objectives, global advantage compares cumulative molecular scores across trajectories, while local advantage compares molecular scores within each trajectory. Adding weighted, centered local advantage distinguishes turns by molecular score while preserving global advantage as their mean turn advantage. Our analysis characterizes the local gradient contribution and gives sufficient conditions for improving the best feasible score over a global-only update. On S2-Bench, L2E improves molar refractivity in 69.7% of cases versus 45.3% for the strongest outcome-level RL baseline under the same proposal budget, while mean relative improvement reaches 1.66×RLOO’s. On MuMOInstruct, gains in joint BBBP, DRD2, and penalized LogP optimization extend to unseen instructions. L2E also benefits from additional refinement turns.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.