Compressing Protein Language Models: A Unified Benchmark
Abstract
Protein language models are becoming central tools across a growing range of tasks, and their scale and adoption are increasing in tandem, making efficiency a main concern, with compression as a popular way to achieve it. Yet what it actually costs remains unmeasured. We present the first compression bench- mark spanning model families, training objectives, and compression paradigms at a common protocol, evaluated at reconstruction, representation geometry, and downstream biology. For deployment, calibration-aware 4-bit is near-lossless at every scale, equivalence-tested with pre-declared margins; 3-bit is free only at the top of a family; and 2-bit collapses the zero-shot channel except at the very top of a family. How much compression costs depends on the readout: a probe retrained on the compressed features survives to 2-bit at scale, whereas a reused head or SAE does not. Underlying this is a structural dissociation: compression largely preserves the representation and shifts its coordinates, so a label-free realignment restores coarse readouts like the probe and SAE, each with a different correction, while the fine-grained LM head reads exactly the part no such map recovers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.