Environment-Conditioned Molecular Spectral Prediction: A Controlled Study of Data, Representation, and Architecture
Abstract
Molecular absorption cross sections govern atmospheric retrieval and biosignature detection, but experimental coverage extends to only a few hundred species, motivating learned emulators. Unlike the single-condition spectra targeted by prior ML spectroscopy work, atmospheric cross sections are environment-conditioned: the same molecule absorbs differently across temperature and pressure. The available labels also carry source-dependent floors. Experimental spectra have an instrumental detection limit near cm/molecule, whereas simulated spectra extend down to . We formulate this task as learning a molecule-conditioned continuous function over wavenumber and evaluate it on 570 experimental HITRAN species under 4-fold molecule-wise cross-validation, comparing 29 configurations across 6 design axes with record-level paired significance testing on a fixed set of 1493 held-out spectra. Synthetic data augmentation is the biggest contributor to test performance on real experimental data: synthetic-only training does not transfer to measured spectra (median ), yet 50% synthetic augmentation in training lifts the median R from 0.46 to 0.60 over experimental training alone. Masking sub-threshold simulated points from the loss beats clipping them at the limit, which is indistinguishable from no treatment. Given these choices, simple encodings win or tie with expressive ones, and equivariant backbones give no advantage over a plain invariant GNN at a matched 5M-parameter budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.