A Counterfactual Framework for Evaluating and Improving the Use of Molecular Structure in Molecular Graph-Language Models
Abstract
Molecular graph-language models have achieved strong performance in molecular description and property prediction. These tasks depend on chemical features encoded in SMILES strings and 2D molecular graphs, including functional groups, ring systems and formal charges. However, current evaluations measure whether an answer is correct, but not whether it is derived from the corresponding chemical structures. Existing datasets also lack controlled molecular pairs that isolate a specific structural feature. In this paper, we propose a counterfactual framework to evaluate and improve how molecular graph-language models use molecular structure. We design a counterfactual benchmark from real, scaffold-matched molecular pairs to test whether models detect the absence of a functional group or another chemical feature. We further design a counterfactual learning model to distinguish molecules with and without the target chemical feature. We also develop an end-to-end training framework that transfers the learned structural information to a broader range of downstream tasks. The benchmark shows that a standard molecular graph-language model achieves only 0.485 pairwise accuracy. Our method raises this accuracy to 0.915 and reaches 0.928 on 822 external molecular pairs. It also improves molecular description, IUPAC naming, molecule-text retrieval, and property prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.