acceptodds
Under review as a conference paper at ICLR 2027

ProtCompass: Biological Information and Predictive Utility in Protein Representations

Abstract

Protein encoders are widely evaluated through structural probing, downstream prediction, and concept erasure, but these measurements can mix learned signal with information already present in the input or introduced by the evaluation itself. We introduce **ProtCompass**, a unified evaluation approach for separating these sources and identifying what protein representations genuinely contribute. ProtCompass audits what each encoder actually reads, compares structural probe scores with non-learned descriptions of the same inputs, evaluates downstream performance against simple sequence baselines, and controls concept-erasure experiments for the representation capacity removed together with the target property. Across diverse protein encoders and biological tasks, we find that structural probe scores can largely reflect retained input information, simple sequence statistics explain substantial downstream performance on several tasks, and many apparent task–property dependencies weaken once concept-erasure effects are properly controlled. These results provide a clearer way to determine what protein representations learn, what they add to biological prediction, and which properties downstream models actually use.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.