Composing Protein Evidence: A Benchmark for Fine-Grained Protein-Text Understanding
Abstract
Protein-text understanding requires more than assigning functional labels: a biological claim is most informative when it is associated with the different levels of evidence. Current protein understanding evaluations often fragment this relation, measuring single global label, local feature, or protein-text matching without testing whether these elements remain bound to the same protein. We introduce PannotGround, a benchmark that turns curated protein annotations into a diagnostic test of evidence composition. PannotGround projects each protein record into synchronized levels and composes them into protein-to-text and text-to-protein tasks with biologically structured hard negatives. Across protein language models, protein-text alignment models, protein-LLMs, and text LLMs, we find that single-level performance does not reliably transfer to compositional binding. Alignment models are strongest, yet remain sensitive to local-evidence contrasts, and many failures are biologically plausible near misses. PannotGround establishes evidence composition as a key diagnostic axis for fine-grained protein-text models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.