Capability-Preserving Molecular Adaptation via Reciprocal Preference Optimization
Abstract
Molecular large language models (LLMs) are commonly adapted from general-purpose LLMs through large-scale domain-specific supervised fine-tuning (SFT). While effective at improving targeted molecular tasks, such specialization may overwrite capabilities already present in the base model, leading to degradation in general-purpose abilities and molecular reasoning capabilities outside the fine-tuning distribution. In this work, we investigate whether preference optimization can provide a more capability-preserving alternative for molecular adaptation. Our key hypothesis is that, by learning relative preferences rather than directly maximizing the likelihood of target responses, while remaining anchored to a reference policy, preference optimization can induce a more conservative adaptation regime than conventional SFT. To make preference optimization feasible without human-annotated molecular preference data, we introduce ReMPA, a Reciprocal Molecular Preference Adaptation framework. ReMPA repurposes existing paired molecule-response supervision into reciprocal preference pairs by pairing template-aligned, structurally similar molecules, such that each response is preferred under its corresponding molecular condition and rejected under the other. Across molecular property prediction, molecular reasoning and comprehension, and evaluations of general capabilities, our experiments demonstrate that, compared with supervised fine-tuning, ReMPA achieves meaningful gains in targeted molecular capabilities while better preserving capabilities already present in the base model. These results suggest that reciprocal preference optimization offers a promising capability-preserving alternative to conventional supervised domain adaptation for molecular LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.