Learning to Align the Known Protein Structure Universe from Sequence Alone
Abstract
Predicted structure databases contain hundreds of millions to over a billion proteins, while sequence collections span tens of billions of nonredundant sequences. How small can a sequence model be while retaining structural alignment quality and compatibility with existing tools? We train ESM-2 encoders through a differentiable affine-gap alignment operator, learning residue scores and gap penalties from TM-align correspondences. On 68,431 SCOPe40-derived validation pairs, fully tuned 8M and 35M models without auxiliary masked language modeling (MLM) achieve full-reference C LDDT of 0.3893 and 0.4185 with maximum expected accuracy (MEA) decoding; a 150M model with MLM reaches 0.3902. All three exceed structure-based and ProstT5-based Foldseek in both 3Di-only and AA+3Di modes. Thresholding correspondence marginals at 0.5 gives F1 scores of 0.4801, 0.5639, and 0.4845 against TM-align, respectively. Frozen-backbone, adapter, LoRA, and initialization controls show that pretraining supplies useful features, but adapting them for alignment matters more than backbone size alone. Scaling beyond 35M offers no benefit under the tested full-tuning recipe with MLM. We then train a head to predict a distribution over Foldseek's 20 3Di states, using their expected score under its fixed substitution matrix. Alignment supervision trains the state assignments without requiring native 3Di labels. At inference, argmax produces pure 3Di states for unchanged Foldseek. In 3Di-only mode, the 8M model reaches full-reference C LDDT 0.3530 against structure-based Foldseek's 0.3506, while the 35M model exceeds both structure-based and ProstT5-based Foldseek on this metric and correspondence F1. These pairwise gains do not transfer uniformly to database retrieval. Compact sequence models can thus learn structural comparisons directly and express them through existing software suites without first predicting coordinates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.