SOUR: Scaling Synthetic Data for Small Molecule Structure Elucidation via Train- and Test-Time Guidance
Abstract
Synthetic data can reduce reliance on scarce experimental annotations, but naively increasing its volume does not guarantee better inverse prediction. This is pronounced for structural elucidation from tandem mass spectra, where high-quality experimental annotation is limited while the forward simulators remain imperfect. We introduce SOUR, a framework for scaling synthetic supervision at training and/or test time by tailoring it to experimental evidence and accommodating imperfect simulations. We use forward simulation to select synthetic examples relevant to experimental spectra of interest and construct virtual refinement trajectories. In training, this supervision supports spectral encoder adaptation and teaches an autoregressive generator to interpret imperfect fingerprints and revise candidate structures. At test time, observed spectra guide query-specific adaptation, while comparisons with newly simulated spectra provide feedback for iterative correction. Our empirical study supports a general strategy for synthetic-data scaling whose benefits transfer across model architectures. On public benchmarks for mass spectrum-based structure elucidation, SOUR-full achieves state-of-the-art identification accuracy, substantially outperforming prior models. These results highlight the potential of imperfect forward simulation to turn an unlabeled observation of unknowns into synthetic supervision at training time and to guide adaptation and correction at test time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.