acceptodds
Under review as a conference paper at ICLR 2027

Scaling Worsens Representation Collapse in Intrinsically Disordered Regions: Pan‑IDR, Evolution‑Guided Contrastive Learning and Ligand‑Dominant Shortcuts for Zero‑Shot Pan‑Cancer IDR‑Ligand Retrieval

Abstract

Learning discriminative representations of intrinsically disordered regions (IDRs) remains a fundamental challenge for protein representation learning. Although IDR targets are commonly assumed to benefit from protein-aware representations, their low-complexity sequences may fundamentally limit the discriminative information available to protein language models. Here, we identify systematic representation collapse across amino-acid composition, ProtBERT, and ESM-2 embeddings on 543 pan-cancer IDR targets. Surprisingly, scaling does not alleviate this problem: increasing ESM-2 from 35M to 3B parameters raises mean pairwise cosine similarity from 0.870 to 0.944, indicating that larger models capture shared IDR sequence statistics more precisely without learning more discriminative representations. We introduce Pan-IDR, a contrastive learning framework combining evolution-guided aggregation (EvoPool), functional attention, and dispersion regularization. Under 5-fold zero-shot evaluation, Pan-IDR achieves a distant-target hit@1 of 0.339, compared with a 0.200 random baseline. Ablations further show that EvoPool is the only component providing a substantial improvement, suggesting that downstream objectives alone cannot recover discriminative information from collapsed representations. More strikingly, a ligand-only ECFP baseline using no protein information outperforms Pan-IDR, achieving distant-target hit@1 of 0.417 versus 0.339. This exposes a ligand-dominant shortcut in which ligand structural similarity provides a stronger predictive signal than learned protein–ligand compatibility, indicating that current IDR affinity data do not support robust protein-aware zero-shot generalization. Together, our results reveal a scaling-resistant limitation of current protein representation learning for IDRs and motivate representations and benchmarks specifically designed for disordered proteins.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.