acceptodds
Under review as a conference paper at ICLR 2027

ProtoCIL: Class-Incremental Learning for Protein Representations with Adaptive Prototypes

Abstract

Protein function annotations continually expand as new enzyme classes are discovered, requiring models to incorporate new classes without retraining from scratch. Yet class-incremental learning for protein function prediction remains underexplored, especially under frozen protein encoders with highly heterogeneous embedding geometries. We present the first systematic benchmark for protein class-incremental learning, covering two enzyme datasets, five frozen protein encoders, and eleven representative CIL baselines. Our benchmark reveals that no fixed-granularity approach, whether shared covariance, random projection, or learned subspace, scales reliably across encoders and class counts. We propose ProtoCIL, a training-free prototype method that adaptively learns group-wise covariance structure over frozen protein embeddings for Mahalanobis classification. ProtoCIL derives its target group count from intrinsic rank and sample size, then refines groups incrementally through validation-based splits as new classes arrive. Across all encoder-dataset configurations, ProtoCIL achieves the highest average incremental accuracy, outperforming the strongest baseline by 8.8% on Enzyme Reaction and 18.5% on EC Number, while maintaining average forgetting below 5%. These results demonstrate that adaptive group-wise covariance sharing provides an effective principle for protein class-incremental learning. Code is provided for review and will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.