State-space, attention or hybrid? PRIMA: a family of foundation models for protein-protein interactions.
Abstract
Protein–protein interactions (PPIs) govern essential biological processes, yet the vast majority of protein language models (PLMs) are transformers trained on single sequences. Our goal is to encode the semantics of the interaction: we ask which architectures are best suited to model PPIs, whether the sequence mixer matters, and at what scale it begins to. We present PRIMA, a family of PPI foundation models that share one training recipe and differ only in their sequence mixer: bidirectional state-space layers (PRIMA-M), full FlashAttention (PRIMA-A), and a hybrid that interleaves one attention layer every state-space layers (PRIMA-H). Each is trained at 8M and 150M parameters in two phases, pre-training on single sequences and post-training on concatenated PPI pairs. We evaluate against the ESM2 family and MINT, the state-of-the-art PPI specialist at 813M parameters, and we rebuild SKEMPI into SKEMPI-R to expose homology leakage and models that resort to recall. At 8M the benchmarks cannot separate the mixers. At 150M they separate, and PRIMA-H matches or outperforms models several times larger. On SKEMPI-R it is first on every metric, reaching Pearson against for ESM2 650M ( larger) and for MINT ( larger), with a partition-to-partition spread tighter, and it leads on the rows where models cannot exploit recall. On Gold Standard PPI, the classification benchmark built to remove protein-identity shortcuts, it matches MINT and PRIMA-M. Scaling PRIMA-H to 650M does not improve its ability to encode interaction specific properties, it only buys a better language model: the extra capacity goes into representing each chain rather than the relationship between them, a different approach is necessary. PRIMA mixers are equivalent after pre-training and diverge after post-training, placing attention's deficit in the interaction objective rather than in protein modelling. We find that PRIMA-H needs no positional encoding at all: permutation-equivariant attention layer can use the position encoded by state-space layers, previously demonstrated only for causal models. At 16,384 tokens PRIMA-H 150M encodes a complex in 74ms and 0.8GB against 10.5s and 78GB for MINT, and runs 65,536-token inputs on which MINT runs out of memory on an 80 GB GPU; below 2,000 tokens the state-space stacks are launch-bound and hold no speed advantage. PRIMA shows that a mixed state-space/attention stack produces specialist-level PPI representations at a fraction of the size and training budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.