Label-Free Fine-Tuning: Your Protein Language Model is Implicitly a Reward Model
Abstract
Post-training of protein language models (PLMs) typically requires verifiable labels from simulations or wet-lab experiments, which are costly to acquire in real world. To overcome the supervision bottleneck, this paper investigates the emerging paradigm of label-free fine-tuning (LFT) to improve protein generation quality and fidelity without ground-truth labels. Prior works have explored various proxy rewards for LFT, yet their design suffer from inconsistent performance across sampling schemes and downstream tasks. To this end, we propose pointwise mutual information (PMI) as a principled reward, which measures the difference between the conditional and unconditional log probability of a given PLM. Despite its simplicity, PMI is motivated by Bayes' rule to effectively leverage a well-trained PLM as an implicit reward model, analogous to classifier-free guidance in diffusion models. Extensive experiments also demonstrate that our framework is amenable and transferable to a broad class of PLMs and protein design tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.