RETHINKING PROTEIN SIMILARITY: ONE REPRESENTATION, MANY RELATIONS
Abstract
Protein retrieval depends on the biological question: finding proteins with a similar fold and finding proteins that bind the same ligand can require different rankings of the same candidates. A fixed similarity score cannot adjust its ranking to the requested criterion, limiting retrieval tailored to different biological questions. We introduce ReaPO (Relation-adaptive Protein Operators), which uses a natural-language instruction and a few directed example comparisons to adapt the similarity rule over frozen protein language model representations. Its signed variant, ReaPO-Signed, assigns relation-dependent positive or negative weights to individual feature directions. We evaluate the approach on crossed ranking, which requires opposite correct rankings for the same query and candidates under two relations, and on ordinary protein retrieval using internal and external datasets. With both relation families withheld from readout training, ReaPO-Signed achieves 0.548 joint accuracy, compared with 0.514 for a conditional positive-semidefinite scorer augmented with a learned global signed scalar that can reverse the whole ranking. When trained separately using only a target-ranking objective, ReaPO-Signed improves nDCG@10 over matched condition-independent signed readouts by 0.0236 on Ordinary20 and 0.0309 on external CATH–BioLiP2 retrieval in the Remote setting. These results support adapting comparison rules to biological relations while reusing shared frozen protein representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.