NOT A FITNESS BOTTLENECK: WHY PROTEIN LANGUAGE MODELS FALL SHORT ON DISORDERED PROTEIN ENSEMBLES
Abstract
Protein language models (PLMs) are increasingly used as frozen encoders to predict the conformational properties of intrinsically disordered proteins (IDPs), yet how much ensemble physics evolutionary pretraining captures remains unclear. Here we measure evolutionary sufficiency (ES), the fraction of a directly supervised model's out-of-distribution accuracy that a frozen PLM representation recovers, for five polymer-physics properties computed from simulated ensembles of nearly 10,000 disordered sequences across six kingdoms, with viral sequences held out. We first show that ES is graded rather than binary, as frozen PLMs adapted to disordered regions recover nearly all of the supervised accuracy for chain dimensions but progressively less for asphericity, the Flory exponent, and its prefactor. We then test and reject a fitness-bottleneck explanation for this shortfall, because the embeddings predict these properties several times better than the PLM's own likelihood and encode net charge almost perfectly, whereas mean pooling hides the charge patterning on which the hardest properties depend. Instead, the gap has a readout component, which a position-aware probe partly closes and which grows with model size, and a pretraining-data component, as encoders pretrained on simulated ensembles reach the supervised ceiling while frozen. Once the backbone is fine-tuned, however, evolutionary initialisation matters little. ES thus offers a probe-aware diagnostic for when frozen PLM features suffice for biophysical prediction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.