Distogram-Based Structural Loss: A Plug-In Module for Protein Inverse Folding
Abstract
Protein inverse folding aims to design amino acid sequences that fold into given backbone structures. However, most inverse folding models are trained primarily to recover native sequences, an objective that does not directly assess structural consistency between generated sequences and the target backbone. Here, we introduce DISCO loss, a model-agnostic auxiliary objective that incorporates pairwise distance information from the target backbone into inverse folding training. A jointly trained auxiliary module predicts residue-pair distance distributions using representations derived from inverse folding outputs, supervised by discretized distances computed from the target backbone. Combined with the sequence recovery loss, DISCO loss provides structural supervision without full 3D reconstruction or separate surrogate pretraining. On de novo protein backbones, DISCO loss improves structural self-consistency across inverse folding models with distinct generation mechanisms, with improvements also observed in long-range contact recovery. Additional evaluations show that sequence recovery performance and diversity are largely maintained. These results demonstrate the effectiveness of auxiliary supervision based on target-backbone distances in improving the structural consistency of generated sequences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.