acceptodds
Under review as a conference paper at ICLR 2027

SOUR: Scaling Synthetic Data for Small Molecule Structure Elucidation via Train- and Test-Time Guidance

Abstract

Synthetic data can reduce reliance on scarce experimental annotations, but naively increasing its volume does not guarantee better inverse prediction. This is pronounced for structural elucidation from tandem mass spectra, where high-quality experimental annotation is limited while the forward simulators remain imperfect. We introduce SOUR, a framework for scaling synthetic supervision at training and/or test time by tailoring it to experimental evidence and accommodating imperfect simulations. We use forward simulation to select synthetic examples relevant to experimental spectra of interest and construct virtual refinement trajectories. In training, this supervision supports spectral encoder adaptation and teaches an autoregressive generator to interpret imperfect fingerprints and revise candidate structures. At test time, observed spectra guide query-specific adaptation, while comparisons with newly simulated spectra provide feedback for iterative correction. Our empirical study supports a general strategy for synthetic-data scaling whose benefits transfer across model architectures. On public benchmarks for mass spectrum-based structure elucidation, SOUR-full achieves state-of-the-art identification accuracy, substantially outperforming prior models. These results highlight the potential of imperfect forward simulation to turn an unlabeled observation of unknowns into synthetic supervision at training time and to guide adaptation and correction at test time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.