Provenance-Trace Policy Optimization for Agentic Retrosynthesis Planning
Abstract
Agentic reinforcement learning (RL) offers a promising approach to retrosynthetic planning. Existing methods train LLM agents to select reactions and explore synthesis routes through tool-mediated interaction, using outcome-based rewards to encourage successful planning. However, trajectory-level rewards do not explicitly distinguish how individual search decisions contribute to planning. Equally rewarded trajectories may resolve different precursor subgoals, while candidate discovery can support reaction selections in later turns. To address this challenge, we propose Provenance-Trace Policy Optimization (PTPO), which converts structured search evidence into local supervision for search decisions. Specifically, PTPO aggregates terminal precursor assessments using the AND–OR synthesis structure and computes local advantages by comparing decisions made for the same molecule under identical ordered candidate views. Through provenance tracing, PTPO further propagates these advantages to the earlier turns that revealed the selected candidates. These local advantages complement trajectory-level supervision through separate turn- and trajectory-level PPO objectives, with each advantage channel normalized independently. Credit estimation reuses the sampled trajectories without a learned critic or additional rollouts. Experiments show that a 4B agent trained with PTPO achieves robust planning performance. On both in-distribution and out-of-distribution retrosynthesis benchmarks, PTPO consistently outperforms outcome-only reward RL agents across all evaluated search budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.