GRASP: Graph Representation Learning with Assay Supervision for Molecular Properties
Abstract
Large unlabeled molecular datasets provide an opportunity to learn reusable representations before expensive property measurements are available. Yet molecular pretraining is commonly followed directly by property-specific fine-tuning, leaving open how large but sparsely labelled bioactivity collections can provide an intermediate source of supervision. We introduce GRASP, a 93.5M-parameter molecular graph Transformer trained through a two-stage representation-learning framework: large-scale self-supervised pretraining followed by sparse multi-assay adaptation. We use replaced-token detection (RTD), where a compact generator produces contextual atom-token replacements and a graph-relative discriminator learns from every valid input position, before the encoder is adapted using sparsely observed labels across 642 ChEMBL assays. Across 23 cluster-held-out OpenADMET endpoints, full fine-tuning achieves a mean MAE of 0.3742, compared with 0.4075 for ECFP4-LightGBM and 0.4395 for Chemprop D-MPNN. Full fine-tuning also improves over frozen features, with LoRA falling between the two. In an OpenADMET comparison of replaced-token and masked-token pretraining, the mean difference is small with frozen features and nearly zero after fine-tuning. Together, these results support a staged approach to molecular representation learning: large-scale self-supervision, followed by sparse assay supervision and downstream fine-tuning. This allows abundant unlabeled molecules and limited experimental measurements to contribute at different stages of training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.