What Counts as a False Discovery in Peptide Sequencing?
Abstract
In peptide sequencing, matching observed fragments does not necessarily establish sequence identity, complicating the interpretation of false-discovery estimates. We investigate this problem using isotope measurements, designed peptide libraries and public sequencing and error-estimation workflows. Isotope traces and fragment matches leave some sequence differences unresolved, whereas synthesis constraints distinguish alternatives under explicit assumptions about library composition. The distinction between fragment support and sequence agreement persists after confidence-based selection. Among spectra with identities supported by both synthesis design and database search, a Casanovo selection at a nominal 5% error rate yields 3.31% disagreement under a fragment-coverage criterion but 11.15% under a modified-sequence criterion. Calibrating confidence against these two criteria also produces different sequence-disagreement rates when the fitted models are applied to other biological studies. Even with the error criterion fixed, observed disagreement depends on the population used for calibration and evaluation, and on whether spectrum matches or distinct peptides are counted. Together, these results show why error estimates must be tied to the sequence identity being claimed, the evidence supporting its assessment and the predictions ultimately reported.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.