Reliable evaluation for small molecule machine learning
Abstract
Cross-validation (CV) allows us to assess the generalizability of machine learning models for small molecule property prediction. Yet, checking whether evaluation results carry over to the target application domain is non-trivial. We introduce a method for comparing any CV split of small molecules against any chosen domain of application. We find that existing splitting methods produce test sets far from the application domain and are prone to data leakage. Next, we introduce a new splitting method based on Maximum Common Edge Subgraph (MCES) distances, and show that resulting CV splits align much better with the application domain of biomolecules. We evaluate five baseline and nine foundation models on different splits, and find that the strongest baseline consistently outperforms all but four of the foundation models on commonly used experimental datasets. All models show a drop of about 0.06 and 0.1 AUROC for MCES splits, indicating that previous splits substantially overstate real-world performance. Notably, pretraining data of foundation models overlaps substantially with downstream evaluation datasets. In the future, our methods may help to reach reliable evaluations for other graph-based machine learning tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.