acceptodds
Under review as a conference paper at ICLR 2027

ProtSent: Protein Sentence Transformers

Abstract

Protein language models produce per-residue representations rich in evolutionary and structural information. The sequence embeddings obtained by mean-pooling those representations, however, are never trained to place related proteins near one another in the embedding space, so the geometry on which retrieval, clustering and nearest-neighbor transfer depend is left to chance. We present ProtSent, a contrastive post-training framework that supervises a protein language model with four types of biological relations at once: Pfam family co-membership, AlphaFold DB Foldseek structural clusters, STRING interactions, and deep mutational scanning fitness ordering. It uses no structure encoder at training or inference, and every training corpus is filtered against the benchmark test sets at 40% identity and 80% coverage. Post-training makes structural family membership recoverable from a frozen embedding without labels. Clustering 2,207 SCOPe-40 domains at their 917 true families raises the adjusted Rand index of the 35M model from 0.054 to 0.507. Family retrieval gains +0.185, +0.190 and +0.424 Recall@1 over three frozen backbones spanning two independently pretrained model families (95% intervals exclude zero). The result reproduces on the CATH v4.3 midnight-zone benchmark, where Homologous-superfamily accuracy rises from 40.7% to 56.7% at 35M and from 43.3% to 62.7% at 150M; these gains of +16 and +19 points exceed the +12 that a CATH-supervised specialist gains over its own ProtT5 base, even though that base is twenty times larger than our backbones. Against sequence-only alignment the frozen embedding leads at ranking depth, MAP 0.705 against 0.607 for a maximally sensitive profile search, and returns a ranking for every query, whereas that search returns no hit at all for 31% of them. We also measure what the reorganization costs, and find that under a trained linear probe the same embeddings fall below the untuned backbone, and that structure distillation pays the same price against its own matched base. Contrastive post-training improves the neighborhoods of the embedding space at the cost of linear-probe accuracy, and that trade is the choice a practitioner faces. We release models at three sizes, the data, and the training recipe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.