Diffusion MSA Pairformer: Structured Masked Diffusion for Protein Representation and Generation
Abstract
Multiple sequence alignments (MSAs) capture evolutionary constraints on protein structure and function through patterns of conservation and coevolution. However, explicit modeling of MSA compositional structure—namely queries, homologous rows, and aligned columns—is notably lacking in current MSA language models. Relying on uniform token masking, existing approaches fail to fully exploit evolutionary information encoded within MSAs. We introduce Diffusion MSA Pairformer, a unified generative model that decomposes MSA diffusion into biologically motivated masking tasks and captures evolutionary dependencies through coupled MSA-pair representations. Empirically, this structured multi-task approach significantly enhances representation learning, with ablation studies confirming its crucial advantage over uniform masking. Our robust approach also enhances generation capabilities: in MSA augmentation and family-conditioned enzyme design, the model achieves superior performance by precisely capturing intra- and interchain structural constraints while balancing sequence diversity, structural consistency, and functional accuracy. These results establish structured MSA diffusion as a principled paradigm unifying evolutionary representation and protein design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.