acceptodds
Under review as a conference paper at ICLR 2027

Label-Free Fine-Tuning: Your Protein Language Model is Implicitly a Reward Model

Abstract

Post-training of protein language models (PLMs) typically requires verifiable labels from simulations or wet-lab experiments, which are costly to acquire in real world. To overcome the supervision bottleneck, this paper investigates the emerging paradigm of label-free fine-tuning (LFT) to improve protein generation quality and fidelity without ground-truth labels. Prior works have explored various proxy rewards for LFT, yet their design suffer from inconsistent performance across sampling schemes and downstream tasks. To this end, we propose pointwise mutual information (PMI) as a principled reward, which measures the difference between the conditional and unconditional log probability of a given PLM. Despite its simplicity, PMI is motivated by Bayes' rule to effectively leverage a well-trained PLM as an implicit reward model, analogous to classifier-free guidance in diffusion models. Extensive experiments also demonstrate that our framework is amenable and transferable to a broad class of PLMs and protein design tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.