acceptodds
Under review as a conference paper at ICLR 2027

AbEvolve: Aligning Protein Language Models with Preference Signals from Antibody Natural Selection

Abstract

While protein language models (PLMs) extract general biological syntax from sequence databases, they often exhibit an alignment gap in antibody engineering. This gap stems from antibodies' unique evolutionary search for optimal binding specificity—a logic that self-supervised PLMs trained on static sequence collections do not explicitly model, because the experimentally labeled data that would reveal this search directly are scarce and costly to produce. In this paper, we instead explore the large-scale, coarse redundancy signal that is readily available in unlabeled antibody sequencing data. To bridge this gap, we propose AbEvolve, a systematic framework that transforms the clonal abundance within antibody sequencing data—a coarse yet massive proxy for natural selection—into explicit abundance-derived evolutionary trajectories for PLM alignment training. By leveraging abundance as a proxy for natural selection, AbEvolve converts latent sequence distributions into actionable signals that help PLMs better capture maturation-associated sequence preferences. To absorb these signals, we adopt two alignment objectives—reward-weighted maximum-likelihood (RW-MLE), which reweights terminal variants by their abundance reward, and GFlowNet alignment, which propagates the scalar abundance rewards along constructed mutation paths and preserves intermediate path structure. We integrate AbEvolve with pre-trained models including ESM2 and ProteinReasoner and evaluate their zero-shot performance across representative antibody property prediction tasks. AbEvolve improves zero-shot ranking on all eight benchmarks—spanning binding, thermostability, and expression—for ESM2-GFlowNet; for ProteinReasoner, GFlowNet alignment improves binding and thermostability, whereas expression gains vary by benchmark. Together, AbEvolve offers a new supervision source and modeling pathway for addressing this alignment gap in therapeutic antibody design.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.