Lost in Sequence, Found in Text: Retrieving Functionally Related Proteins Under Joint Sequence and Structure Similarity Controls
Abstract
Convergent evolution can give rise to proteins with similar functions despite dissimilar sequences and structures, motivating functional retrieval beyond conventional similarity search. Yet the ability of existing methods to recover such functional relationships remains insufficiently evaluated under joint controls on sequence and structural similarity. We introduce Fureton, a FUnctional RETrieval benchmark for this setting built from 89,116 reviewed UniProtKB/Swiss-Prot proteins with AlphaFold-predicted structures and selected Gene Ontology (GO) annotations. Designed to assess functional retrieval beyond conventional similarity, the benchmark controls both sequence and structural similarity between query–target pairs and defines relevance through informative shared GO ancestors. We evaluate five families of methods in zero-shot settings: conventional sequence-, structure-, and signature-based search, protein sequence language models, protein structure-aware language models, protein–text contrastive models, and protein-to-text generative models. We additionally adapt protein language models by training projection heads with a contrastive objective while keeping the pretrained backbones frozen. This contrastive adaptation improves retrieval across Molecular Function, Biological Process, Cellular Component, and the overall task, suggesting that it can make pretrained protein representations more effective for functional retrieval. Nevertheless, the best adapted model achieves 4.65% MAP in overall task, compared with 7.65% for retrieval using protein-to-text generative models. These results highlight functional text as a promising representation for retrieving functionally related proteins under joint sequence and structural similarity controls. Code, data, and evaluation protocols are available at https://anonymous.4open.science/r/FURETON-EA8D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.