A Multi-Axis Out-of-Distribution Benchmark for Protein Understanding Tasks
Abstract
Protein representation learning has advanced rapidly, driven by large-scale language models and structure encoders. However, most existing benchmarks evaluate models under independent and identically distributed (IID) settings or rely on a single out-of-distribution (OOD) axis, such as sequence identity. In real-world applications, test proteins frequently diverge from the training distribution along multiple biological dimensions—sequence, structure and function—yet current benchmarks do not systematically measure model robustness under these diverse shifts. To address this gap, we introduce POOR (Protein Out-Of-distribution Robustness), a multi-axis OOD benchmark for protein understanding that spans 10 protein tasks, each annotated with 24–28 OOD scenario flags covering sequence/evolutionary OOD, structural OOD, and functional OOD. We further propose ProDPR (Protein Distributional Performance Retention), a unified OOD evaluation metric that enables comparison of model generalization difficulty across both single-axis OOD scenarios and their high-dimensional combinations. Using this framework, we benchmark protein language models at multiple scales and structure-aware models, and further examine the effects of different fine-tuning regimes. We find that stronger in-distribution performance does not necessarily imply greater OOD robustness: model scale and structural inductive bias confer distinct, shift-dependent advantages, while fused multi-axis shifts expose substantially larger failures than conventional single-axis evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.