Error-Controlled Discovery of Protein Interactions via Conformal Calibration
Abstract
Protein language models (PLMs) are powerful predictors of protein interactions, but strong benchmark performance does not ensure reliable discoveries. We introduce PPI-Caliper, a model-agnostic framework that adapts conformal calibration to select interactions at user-specified false discovery rate (FDR) levels using held-out labels without retraining the predictor. Across five species, two PLM-based predictors, and two human training-data versions for each predictor, PPI-Caliper keeps empirical FDR at or below the target level under in-distribution calibration. Under distribution shift, transferred thresholds tend to become conservative when predictive performance improves, but can fail to control FDR when performance deteriorates. PPI-Caliper connects pre-trained interaction scores to explicit error targets and demonstrates why reliable biological discovery requires evaluating calibration alongside predictive accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.