acceptodds
Under review as a conference paper at ICLR 2027

Constructed Negatives Shape the Evaluation of Sequence-Based Protein–Protein Interaction Models

Abstract

Many machine-learning tasks, including representation learning, rely on negative examples that are constructed rather than directly observed. This issue is especially important for protein–protein interaction (PPI) prediction, where experimentally established non-interactions are scarce. PPI models are therefore commonly trained and evaluated against sampled, unobserved protein pairs treated as negatives. We study how this construction affects model comparison. Keeping positives, protein-disjoint splits, class balance, and model selection fixed, we vary training and evaluation negatives independently across six constructions and four sequence-based models, three of which use protein language model (pLM) representations. Changing only the test negatives shifts AUPRC by up to 0.26, and the top-ranked architecture changes with the construction. Negatives matched to true partners in pLM representation space reduce all pLM-based models to near-chance performance, and shuffling partners within each class shows that the original pairing contributes little to AUROC. Constructed training negatives rarely outperform uniform sampling: on external data, they underperform it in 108 of 120 comparisons for the pLM-based models, with gains never exceeding 0.05 AUROC and losses up to 0.23. We release splits, negative-generation procedures, and evaluation code that allow training and evaluation negatives to be varied separately. Our results show that with constructed negatives, model comparisons are not just noisy but systematically sensitive to how negatives are defined, a methodological choice that is often left implicit.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.