acceptodds
Under review as a conference paper at ICLR 2027

Physio-DPO: Aligning Protein Language Models with Local Energy Gaps to Mitigate Structural Hallucinations

Abstract

Protein language models generate sequences that score well under their own likelihood yet fold into structures an empirical energy function rates as unstable. Direct Preference Optimization cannot separate such a case from a marginal one, because a binary label discards how large the violation is. We propose Physio-DPO, which recovers that magnitude. For a backbone-matched pair consisting of a native structure and a physics-perturbed decoy of the same fold, a thresholded and bounded function of the local energy gap scales the DPO update, so severe stereochemical violations drive larger updates than ambiguous ones. Pairs of unrelated folds keep the ordinary binary objective, so no energy magnitude is ever compared across folds. We build PhysioPref-1M, one million preference pairs that combine generation-mined preferences with physics-perturbed hard negatives, and verify the labels by blinded expert review. Across three seeds Physio-DPO reaches an sc-RMSD of  \AA and a foldability of , ahead of SFT, PPO, DPO, and the strongest energy-aware baseline, and it keeps that lead under folding and scoring oracles unseen during training. The gains reach experimental data as well: on ProteinGym deep mutational scanning assays Physio-DPO raises the average Spearman correlation from to .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.