Tabular In-Context Learning for Structural Generalization in Molecular Property Prediction
Abstract
Molecular property prediction helps identify compounds with desired characteristics by estimating molecular properties from chemical structures. In applications such as drug discovery and materials design, this often requires generalizing to molecules that are structurally different from those in the labeled training data. Recent studies have applied tabular foundation models (TFMs) to molecular property prediction through in-context learning, but their effectiveness under such structural train–test shifts remains underexplored. To address this gap, we evaluate TFM-based approaches in this setting and introduce MolRIFT (MOLecular Relational Inference with Frozen TFMs), a framework that incorporates explicit molecular comparisons into tabular in-context learning to improve structural generalization in molecular property prediction. In MolRIFT, one TFM makes an initial prediction from molecule-level context, while a second TFM uses molecular pairs and their prediction-error differences to refine it. Across 58 MoleculeACE and Polaris tasks, combining CheMeleon representations with TabPFN-3 wins a majority of tasks against each evaluated baseline, while the MolRIFT framework improves this strong predictor on 46 of the 58 tasks. The gains also persist across four molecular representations and three TFM backbones, supporting MolRIFT as a broadly applicable approach. The code and datasets are available at https://anonymous.4open.science/r/MolRIFT-8773/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.