POSEIDON: Reinforcement Learning for Molecular Design with LLM-Guided Structural Feedback
Abstract
Reinforcement learning (RL) is a powerful approach to goal-directed molecular generation, yet docking-driven workflows commonly reduce protein–ligand evaluations to scalar rewards. This compression leaves rich information in docked poses underutilized, including spatial complementarity, interaction patterns, and local structural deficiencies that could guide molecular improvement. We propose POSEIDON, a framework that integrates a frozen large language model (LLM) into the learning loop of an RL-based chemical language model (CLM). POSEIDON translates docking-derived structural evidence into explicit, interpretable molecular edits, re-evaluates the resulting candidates through docking, and incorporates verified outcomes into a shared candidate memory. These candidates are replayed to update the CLM, allowing structure-guided improvements to influence subsequent molecular generation. To make this feedback loop adaptive and cost-efficient, POSEIDON combines three components: an evolving domain-specific prompt informed by verified editing outcomes, a budget-aware allocation strategy that prioritizes promising opportunities for LLM intervention, and a source-aware memory and replay mechanism that balances candidate quality, chemical diversity, and exploration. Together, these components connect structural reasoning, targeted editing, and policy learning within a unified optimization workflow. We evaluate POSEIDON on the CrossDocked2020 benchmark and a bile acid molecular design case study, examining optimization efficiency, molecular diversity, and adherence to structural constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.