acceptodds
Under review as a conference paper at ICLR 2027

Self-Conditioned Denoising for Atomistic Representation Learning

Abstract

Large-scale pre-training has transformed NLP and computer vision and is motivating analogous foundation models for the physical sciences, yet label free pre-training strategies for atomistic data remain underexplored. To date, supervised pre-training on DFT force-energy labels has given the largest gains in downstream property prediction, outperforming self-supervised learning (SSL) methods that remain limited to ground-state geometries or single domains of atomistic data. We address these shortcomings with Self-Conditioned Denoising (SCD), a backbone-agnostic reconstruction objective in which the model's own embedding of a clean geometry, its self-embedding, conditions the denoising of a corrupted copy. SCD applies to any atomistic domain, including small molecules, proteins, periodic materials, and non-equilibrium geometries. With backbone and pre-training data matched, SCD reduces error by 20–45% relative to standard node denoising on QM9 targets, outperforms the Frad and SliDe variants on most targets, and matches the performance of supervised force-energy pre-training on identical geometries. A small, fast GNN pre-trained by SCD is competitive with or superior to larger models pre-trained on far larger labeled or unlabeled datasets across multiple domains. We also fine-tune leading foundation potentials (MACE-OFF, Orb-v3, UMA) for property prediction; SCD models remain competitive with them at one to two orders of magnitude less pre-training compute.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.