Simple Models Can Still Compete in Out-of-Distribution Molecular Property Prediction
Abstract
This study began as a replication of BOOM, a recent benchmark for out-of-distribution (OOD) molecular property prediction (Antoniuk et al., 2025), and developed into a broader investigation of how model choice, evaluation design, and metrics affect conclusions about molecular OOD performance. We reproduce and extend BOOM using six models spanning linear, tree-based, message-passing, pretrained language, and 3D-equivariant approaches. A simple Elastic Net model based on 2D molecular descriptors performs competitively with Chemprop on the original property-based OOD tasks, showing that increased model complexity does not necessarily translate into better extrapolation. This conclusion remains qualitatively robust across repeated runs and after validation-based hyperparameter tuning of Chemprop. In contrast, Random Forest and XGBoost perform poorly on these tasks, illustrating that the ability to extrapolate beyond the training target range, rather than model complexity alone, strongly influences performance. We further show that BOOM's original OOD evaluation measures extrapolation to extreme property values rather than structural novelty. When we introduce structure-based OOD splits, all evaluated models perform substantially better and differences between them narrow. Finally, we identify a metric discrepancy in the original benchmark: values reported as correspond to squared Pearson correlation (). Using the coefficient of determination instead reveals many strongly negative OOD scores and a more severe extrapolation problem. Together, these results show that simple models remain important baselines for molecular property prediction and that conclusions about OOD performance depend critically on the type of distribution shift, the models included, and the evaluation metric.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.