Physio-DPO: Aligning Protein Language Models with Local Energy Gaps to Mitigate Structural Hallucinations
Abstract
Protein language models generate sequences that score well under their own likelihood yet fold into structures an empirical energy function rates as unstable. Direct Preference Optimization cannot separate such a case from a marginal one, because a binary label discards how large the violation is. We propose Physio-DPO, which recovers that magnitude. For a backbone-matched pair consisting of a native structure and a physics-perturbed decoy of the same fold, a thresholded and bounded function of the local energy gap scales the DPO update, so severe stereochemical violations drive larger updates than ambiguous ones. Pairs of unrelated folds keep the ordinary binary objective, so no energy magnitude is ever compared across folds. We build PhysioPref-1M, one million preference pairs that combine generation-mined preferences with physics-perturbed hard negatives, and verify the labels by blinded expert review. Across three seeds Physio-DPO reaches an sc-RMSD of \AA and a foldability of , ahead of SFT, PPO, DPO, and the strongest energy-aware baseline, and it keeps that lead under folding and scoring oracles unseen during training. The gains reach experimental data as well: on ProteinGym deep mutational scanning assays Physio-DPO raises the average Spearman correlation from to .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.