PROGEA: Unifying Protein Sequence, Structure, and Language with Diffusion Language Models
Abstract
We envision a general-purpose protein foundation model as a unified, interactive model of sequence, structure, and function, capable of understanding, generation, and cross-modal translation while retaining the instruction-following and reasoning capabilities of modern LLMs. Toward this vision, we introduce **PROGEA**, the first native, unified, multimodal diffusion language model (dLLM) for proteins. **PROGEA** is built on two key design choices: (1) a unified discrete token space that jointly represents sequence, structure, and natural language; and (2) a dLLM backbone that replaces strict left-to-right factorization with bidirectional or semi-bidirectional token interactions, enabling global residue communication and long-range contact modeling that are naturally suited to structure modeling. This formulation unifies six core capabilities: folding, inverse folding, sequence understanding, structure understanding, sequence design, and structure design, while retaining the native language capabilities of general-purpose LLMs. On established benchmarks, **PROGEA** achieves broadly competitive performance. In folding, PROGEA-MDLM improves backbone TM-score over existing tri-modal models by 0.30/0.16 on CAMEO22/PDB split. For sequence understanding, PROGEA-BDLM closely matches the best evaluated frontier LLM, Claude Opus 4.7 (62.37%). These results provide strong evidence for effective sequence–structure alignment, extending the sequence–language representation space established in prior work into the 3D physical space of proteins.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.