NO PROTEIN LANGUAGE MODEL BEATS A LOOKUP TABLE AT FOLD-LEVEL RETRIEVAL
Abstract
Remote homology search finds proteins that share a fold after their sequences have diverged past the range where alignment scores work. The field solves it by embedding every protein with a pretrained protein language model. However, whether those embeddings are what carries fold-level ranking has not been tested across the released models, because published comparisons are small and leave the readout layer, label leakage and index size uncontrolled. Here we evaluate 115 released protein language models from 2.6M to 16.1B parameters on SCOPe40 retrieval under one protocol. We compare them with the Additive Normalised Gated K-mer index (ANGK), a static index that answers a query by table lookup with no forward pass. Fold-level AUROC falls as parameter count grows, with a Spearman correlation of −0.57 across the 115 models. The 39.9 MiB index reaches 0.9555 where no model reaches 0.9500. It also encodes eleven times faster than the fastest of them and holds first place at the topology level of a second taxonomy. A four-way position bin on 3-mers supplies that margin. The same index covers 2.41 × 108 AlphaFold structures in 7.4 hours. One pass over them finds 4,931 structures that connect 53 fold pairs SCOPe keeps apart. Fold-level retrieval now has a baseline with no forward pass.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.