PROPEL: Global Evolutionary Latents for Autoregressive Protein Language Models
Abstract
Autoregressive protein language models capture rich sequence regularities, but represent protein family information implicitly through shared model parameters and token-level hidden states. Other family-specific models about one family are spread across weights shared with every other family, so pointing it at a new target means changing the model. We introduce PROPEL (Protein Retargeting by Optimized Posterior over Evolutionary Latents), an autoregressive protein language model with latent space that complements local autoregressive modeling with a compact global representation of evolutionary context. A target-specific latent state is shared across the sequence and injected into every decoder layer through cross-attention, separating persistent family-level information from token-level sequence computation.The latent is learned without an encoder and can be further optimized at inference time while the backbone remains frozen. This representation provides a flexible interface for incorporating both unlabelled homologs and limited fitness supervision, while preserving a single autoregressive model for both variant scoring and sequence generation. Across protein fitness prediction tasks, PROPEL consistently improves its autoregressive backbone and exhibits strong data and parameter efficiency: a 60M-parameter model can match substantially larger autoregressive protein language models under zero-shot fitness prediction and comparable alignment settings. These results show that introducing an explicit global representation can substantially enhance autoregressive protein models without sacrificing their simplicity, scalability, or generative nature.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.