Beyond Static Annotation Matching: A Temporal Benchmark for Free-Form Protein Function Prediction
Abstract
Advances in large language models have enabled protein models to generate detailed natural-language descriptions of protein function. However, evaluating such free-form predictions remains challenging, as existing benchmarks are primarily built around fixed labels or static annotations. We introduce ProLong-Bench, a temporal benchmark for free-form protein function prediction that evaluates models as protein annotations develop over time. Across six representative Protein LLM systems, we observe substantial differences in performance across temporal settings. Models that perform well on existing functional annotations do not necessarily capture knowledge added later, while performance on newly added proteins shows a different pattern. These results highlight temporal evaluation as an important complement to conventional static benchmarking. We intend ProLong-Bench to provide a standardized and extensible evaluation framework for free-form protein function prediction, facilitating more comprehensive model comparison and supporting future progress in the field.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.