Multi-Peptide Prompting Enables In-Context Learning in Protein Language Models
Abstract
We show that single-sequence PLMs can perform in-context peptide learning without gradient updates, task-specific retraining, or architectural modification. We introduce multi-peptide example prompts (MPEPs), in which demonstration peptides are concatenated with glycine spacers and used as context for scoring query peptides by their prompted probability. We evaluate this approach across multiple tasks, using both encoder-only ESM-2 models and decoder-only ProGen2 models. Across most tasks, performance improves with the number of peptide examples and with model scale, indicating that PLMs can extract shared properties from prompted examples. We further introduce a difference score that contrasts positive-example and negative-example MPEPs, reducing compositional biases in raw PLM probabilities and substantially improving classification. Perhaps surprisingly, when applied to few-shot peptide learning problems MPEP-based classification is competitive with classifiers trained on frozen ESM-2 embeddings, despite requiring no training. These results reveal an unexpected in-context inference capability in single-sequence PLMs and establish in-context learning as a potential lightweight strategy for low-data peptide classification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.