MolPFN: Graph-Level Prior-Data Fitted Networks for Molecular Property Prediction
Abstract
Molecular property prediction is essential across chemical, biological, and materials applications, from molecular discovery to materials design. However, experimentally measured labels are often scarce, whereas existing molecular models often require large labeled datasets and task-specific training. Tabular foundation models based on Prior-Data Fitted Networks (PFNs) offer a promising alternative: they achieve strong predictive performance in small-data regimes and can adapt to new tasks through in-context learning without task-specific parameter updates. Yet, applying these models to molecules requires bridging the gap between graphstructured molecular inputs and the tabular representations expected by PFNs. We introduce MOLPFN, a framework for in-context molecular property prediction with a learnable molecular graph adapter, that transforms molecular graphs into PFN-compatible representations, via a synthetic graph prior designed for PFN pretraining. MOLPFN thereby combines molecular graph structure and atomic information with the in-context learning capabilities of PFNs, while supporting optional auxiliary molecule-level information, such as learned embeddings and molecular fingerprints, without task-specific retraining. We evaluate MOLPFN on nine molecular benchmarks covering regression and classification tasks. When trained only on tasks sampled from the synthetic graph prior, MOLPFN matches or outperforms established tabular foundation models consistently improving results across tasks. We further fine-tune MOLPFN across thousands of experimentally measured small-molecule bioactivity tasks, substantially improving fewshot prediction for previously unseen protein targets and outperforming most established few-shot molecular learning baselines, even when using only standard atomic features. Together, these results show how domain-specific adapters and structured priors can extend tabular foundation models to molecular graphs, enabling flexible in-context prediction across molecular tasks and representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.