Multi-Teacher On-Policy Distillation for Multi-Property Molecular Optimization
Abstract
Lead optimization requires a generative policy to improve every property named in an instruction while keeping the molecule valid and structurally close to the lead. Property specialists trained with reinforcement learning are cheap to obtain, but composing them is not: a molecular instruction activates several specialists at once, they differ in how dependable they are at a given molecular state, and specialists that propose comparably reasonable molecules may still prefer different continuations of the same partial structure. We propose reliability-aware multi-teacher on-policy distillation (RA-MOPD), which folds a bank of frozen property experts into a single policy that queries neither experts nor property oracles at inference. At every state the student visits, RA-MOPD weights the requested experts by held-out reliability together with verified outcomes of their own proposals at that state. The weighted experts are combined into a full-vocabulary product of experts on the student's realized prefixes, and we show that the log-normalizer of this product is exactly the minimum weighted reverse-KL disagreement among the active experts, which we use as a token-local gate on the distillation loss. A group-relative verifier reward supplies the complementary trajectory-level signal, scoring complete recurrent edits, while the weighted teachers provide token-level distributional supervision on the molecular answer. Across ten three- and four-property tasks, RA-MOPD improves constrained success over supervised fine-tuning, verifier-only reinforcement learning, offline expert imitation, parameter merging and uniformly weighted multi-teacher distillation, and its advantage widens as every property is required to improve by a larger margin.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.