PAIRank: A MODEL-AGNOSTIC RELIABILITY PLUGIN FOR PROTEIN–LIGAND BINDING AFFINITY PREDICTION
Abstract
tional drug screening, but their performance can deteriorate on novel chemical scaffolds and other out-of-distribution (OOD) samples. Most predictors do not provide sample-level reliability information for identifying predictions with po- tentially large errors. Here, we introduce PAIRank (Protein–ligand Affinity In- dividual Reliability Ranking), a model-agnostic reliability plug-in for protein– ligand binding affinity prediction. PAIRank extracts joint protein–ligand embed- dings from a frozen base affinity predictor, partitions the training embeddings into local regions and computes cluster-conditioned Mahalanobis distances from each evaluated sample to every region. Together, these distances describe how the sam- ple is positioned relative to the training distribution. Pairwise supervision derived from base-model prediction errors then enables PAIRank to learn the relative error risk of different samples and assign a score for reliability ranking. These scores allow lower-risk predictions to be retained preferentially. We evaluated PAIRank under ligand-scaffold extrapolation on ChEMBL and on a temporally external BindingDB set using multiple affinity predictors and uncertainty-quantification methods. Experimental results showed that PAIRank reduced error in the re- tained subsets and exhibited competitive and relatively stable selective-retention performance across base models and data distributions. PAIRank requires neither modification nor retraining of the base predictor and it can serve as a lightweight post-processing plug-in for assessing the reliability of individual predictions from different affinity-prediction frameworks under distribution shift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.