acceptodds
Under review as a conference paper at ICLR 2027

NO PROTEIN LANGUAGE MODEL BEATS A LOOKUP TABLE AT FOLD-LEVEL RETRIEVAL

Abstract

Remote homology search finds proteins that share a fold after their sequences have diverged past the range where alignment scores work. The field solves it by embedding every protein with a pretrained protein language model. However, whether those embeddings are what carries fold-level ranking has not been tested across the released models, because published comparisons are small and leave the readout layer, label leakage and index size uncontrolled. Here we evaluate 115 released protein language models from 2.6M to 16.1B parameters on SCOPe40 retrieval under one protocol. We compare them with the Additive Normalised Gated K-mer index (ANGK), a static index that answers a query by table lookup with no forward pass. Fold-level AUROC falls as parameter count grows, with a Spearman correlation of −0.57 across the 115 models. The 39.9 MiB index reaches 0.9555 where no model reaches 0.9500. It also encodes eleven times faster than the fastest of them and holds first place at the topology level of a second taxonomy. A four-way position bin on 3-mers supplies that margin. The same index covers 2.41 × 108 AlphaFold structures in 7.4 hours. One pass over them finds 4,931 structures that connect 53 fold pairs SCOPe keeps apart. Fold-level retrieval now has a baseline with no forward pass.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.