acceptodds
Under review as a conference paper at ICLR 2027

PROGEA: Unifying Protein Sequence, Structure, and Language with Diffusion Language Models

Abstract

We envision a general-purpose protein foundation model as a unified, interactive model of sequence, structure, and function, capable of understanding, generation, and cross-modal translation while retaining the instruction-following and reasoning capabilities of modern LLMs. Toward this vision, we introduce **PROGEA**, the first native, unified, multimodal diffusion language model (dLLM) for proteins. **PROGEA** is built on two key design choices: (1) a unified discrete token space that jointly represents sequence, structure, and natural language; and (2) a dLLM backbone that replaces strict left-to-right factorization with bidirectional or semi-bidirectional token interactions, enabling global residue communication and long-range contact modeling that are naturally suited to structure modeling. This formulation unifies six core capabilities: folding, inverse folding, sequence understanding, structure understanding, sequence design, and structure design, while retaining the native language capabilities of general-purpose LLMs. On established benchmarks, **PROGEA** achieves broadly competitive performance. In folding, PROGEA-MDLM improves backbone TM-score over existing tri-modal models by 0.30/0.16 on CAMEO22/PDB split. For sequence understanding, PROGEA-BDLM closely matches the best evaluated frontier LLM, Claude Opus 4.7 (62.37%). These results provide strong evidence for effective sequence–structure alignment, extending the sequence–language representation space established in prior work into the 3D physical space of proteins.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.