Rate-Constrained Geometry-Aware Protein Inverse Folding with Synchronized Masking
Abstract
Structure-conditioned protein design requires bridging discrete one-dimensional sequence representations and continuous three-dimensional geometry while preserving evolutionary and structural constraints. We propose GeoSeqAlign, a probabilistic inverse-folding framework that aligns sequence and structure through a variational information bottleneck (VIB). To prevent trivial reliance on local cross-modal correspondences, we introduce synchronized structure–sequence corruption, which masks both modalities at the same residue positions. Combined with a bottleneck that balances reconstruction fidelity against latent information capacity, this mechanism encourages the model to capture predictive sequence–structure relationships while reducing sensitivity to geometric perturbations. To integrate geometric constraints with evolutionary priors, GeoSeqAlign injects -invariant pairwise distances and relative orientations as attention biases into a frozen protein language model, enabling geometry-aware sequence generation without coordinate-equivariant updates. Across CATH+ 4.3, TS50, and CASP15/TS45, GeoSeqAlign achieves the highest amino-acid recovery among the evaluated fixed-backbone baselines while tuning only 6.4% of its parameters. Controlled ablations support the complementary contributions of bottleneck regularization and synchronized corruption, with greater benefits under structural distribution shift and backbone perturbations. These results highlight information-constrained cross-modal learning as an effective approach to robust, parameter-efficient protein inverse folding. Source code is available at: https://anonymous.4open.science/r/GeoSeqAlign-17466.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.