ProtGPT3: An Open-source family of Promptable and Aligned Protein Language Models
Abstract
Generative protein language models (pLMs) enable exploration of vast sequence spaces, but controlling generation toward desired functional families remains challenging, particularly in low-data regimes. We introduce ProtGPT3, an open family of autoregressive pLMs from 112M to 10B parameters, comprising single-sequence and MSA-promptable models and, to our knowledge, the largest publicly released pLM trained exclusively with an autoregressive objective. We use ProtGPT3 to compare supervised fine-tuning against few-shot prompting with homologous sequences. Across 25 held-out protein families, prompting outperforms fine-tuning in the low-data regime and matches it as training-set size grows. In an extreme low-data case study on defluorinase enzymes, we generate enzymes by fine-tuning or prompting ProtGPT3 models with seven enzymes only, producing four new defluorinase with activity values in line with the natural ones. We further investigate two complementary ways to control generation. First, we apply direct preference optimization to align the models at a large scale, mitigating the well-established pre-training bias toward low-complexity sequences, while preserving sequence diversity. Second, we introduce a novel Feynman-Kac sequential Monte Carlo procedure that uses inference-time compute to steer generation toward sequence-level objectives without updating model parameters. We release model weights, configurations, training and inference code on Hugging Face ( https://huggingface.co/protgpt3 ) under a pseudonym.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.