Bridging Structure to Sequence via Local Geometries for Protein Inverse Folding
Abstract
Inverse folding—designing amino acid sequences for a given protein structure—is a core task in protein engineering. Recently, masked diffusion protein language models (PLMs) have shown great potential, typically decoding residues in a confidence-first order. While decoding order has drawn growing attention in diffusion large language models (dLLMs), it has remained largely unexplored in protein language models. In this work, we evaluate decoding orders through a unified sequence–structure pass@K metric, assessing a model's ability to generate sequences that are both native-like and designable. Unlike in natural language dLLMs, where decoding order reshapes sampling coverage, we find that in inverse folding it leaves coverage unchanged but steers a trade-off: confidence-first decoding yields highly native-like but poorly designable sequences. Entropy analyses trace this to premature commitment: high-uncertainty positions are deferred and resolved by sequence statistics rather than structural geometry. We show the trade-off can be addressed at two levels. At inference time, a local-geometry guided decoding order resolves structurally prioritized regions first while sampling freely within them. At training time, upweighting high-uncertainty positions during fine-tuning biases the model toward structural constraints at exactly those positions. On CATH, the resulting model improves designability and sequence recovery simultaneously over its confidence-decoded counterpart, with consistent gains on the out-of-distribution TS50 and TS500 benchmarks and strong performance on motif-scaffolding. This plug-and-play framework enables both inference-time control and training-time enhancement.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.