Iona: Scaling Law for Mass Spectrometry Foundation Model
Abstract
Mass spectrometry is a core analytical technique for identifying chemical molecules, with applications ranging from cancer biomarker and metabolite detection to heavy isotope analysis, pathogen identification, and explosive screening at airports. Building machine learning models that genuinely understand mass spectra is therefore an important problem. Recent transformer models trained on tandem mass spectra span roughly 1M to 400M parameters, yet a basic question remains unanswered: *do larger models actually understand spectra better, and how large is large enough?* No published spectrum model reports a systematic size sweep or a fitted scaling law. In this work, we conduct the **first such study for self-supervised mass spectrometry foundation models**. We present **Iona**, a family of encoders pretrained at 25M, 50M, 100M, 200M, and 400M parameters on MSConsensus-100M, a corpus of 100 million consensus MS/MS spectra, using a masked intensity-reconstruction objective. Evaluation loss follows a power law in pretraining compute, with a compute-efficient frontier on which each larger model overtakes the previous one, and extrapolating the frontier places the compute-optimal model size well beyond the largest spectrum transformer published to date; we therefore aim to push Iona past 1 billion parameters next. Capability grows with scale downstream as well, from denoising and retrieval to chimeric-spectrum decomposition, and with simple fine-tuning Iona reaches or surpasses strong specialised models in real identification pipelines: it beats a supervised proteomics encoder on real chimeric spectra, adds 20,763 peptide–spectrum matches at 1% FDR to database search, and recovers 37–60% more peptides from the weakest single-cell acquisitions. Random-initialisation controls show that this performance comes from pretraining, not fine-tuning. We further test a previously unexamined hypothesis: *can a model succeed without encoding any absolute ?* In our design, all mass information enters through a learned Fourier-feature bias on pairwise mass differences, making the encoder translation-invariant in . Contrary to the assumption that absolute mass positioning is necessary, this relative-only design solves every task above, and its learned bias recovers amino-acid residue, neutral-loss and isotope masses without supervision: the model learns chemical reality from data. Because the recipe uses no peptide labels, annotations or any other proteomics-specific ingredient, we expect the same scaling behaviour to hold for metabolomics and other mass spectrometry applications.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.