FreeDIA: Search-Free Pretraining for DIA De Novo Peptide Sequencing
Abstract
De novo peptide sequencing reconstructs peptide sequences directly from tandem mass spectra without relying on a reference database search. However, sequencing from data-independent acquisition (DIA) spectra is challenging because peaks from multiple co-isolated precursors and background signals are mixed within the same spectrum. Existing DIA de novo methods often use search-dependent supervision based on database-identified peptides, limiting the use of unidentified DIA data for representation learning. We introduce FreeDIA, a search-free pretraining framework that learns precursor–fragment attribution from sequence-independent chromatographic supervision. Instead of using rule-based assignments as hard filters, FreeDIA uses them as supervision while retaining the complete observed spectrum, jointly learning denoising and deconvolution and transferring precursor-specific fragment evidence to peptide decoding. Across two datasets, FreeDIA achieves strong performance under both unseen-peptide and sample-holdout evaluations, with state-of-the-art generalization to unseen peptide sequences. Further analyses show that the learned attribution captures precursor-specific and graded spectral evidence beyond binary assignments, while ablation studies confirm the complementary roles of pretraining and reranking. These results highlight search-free spectral pretraining as an effective way to exploit DIA data without requiring peptide identifications during pretraining. Code for reproducing our experiments is provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.