acceptodds
Under review as a conference paper at ICLR 2027

SkeleMol: Molecular Property Prediction with Skeletal Images

Abstract

Molecular property prediction is central to drug discovery and chemistry, as it lets candidate molecules be screened before synthesis and testing. It typically relies on specialized representations such as molecular graphs, SMILES strings, or 3D conformers, together with models pre-trained specifically on molecular data. Models for each representation are growing increasingly complex, and the field is fragmenting along input modalities. We investigate a different approach: representing molecules as skeletal images and adapting general-purpose vision encoders to molecular prediction. We introduce , an image-based framework that uses only a molecular drawing at inference time. It enables pre-trained vision encoders for molecular property prediction, replacing the molecule-specific architecture. During pre-training, we progressively inject chemical and geometric information through image-SMILES alignment, distillation from a 3D molecular model, and a chemistry-informed curriculum. Across 10 vision architectures and seven pre-training strategies, we find that the choice of molecular supervision has a substantially larger effect than the choice of vision backbone, with consistent strategy rankings across architectures. In particular, transferring 3D knowledge improves an image encoder despite the absence of conformers at inference, while curriculum learning further improves data efficiency. Across 10 benchmarks spanning physical, biological, and quantum properties, achieves state-of-the-art results on five tasks using 2 M pre-training molecules, while retaining a standard vision architecture and image-only inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.