acceptodds
Under review as a conference paper at ICLR 2027

Designing Protein Structures from Natural Language

Abstract

A central challenge in de novo protein design is generating proteins with desired structural and functional properties. Existing design models typically require a dedicated input representation for each property, limiting the design objectives they support. We develop Prompteina, the first all-atom structure generation model that supports new and broader design objectives directly specified in natural language, analogous to text-to-image generation. To train Prompteina, we construct Prompteina-Data, a paired text–protein structure dataset of 1.8M structures and 50M+ captions describing global and residue-specific structural and functional properties. We then adapt Proteina-Complexa, a pretrained structure generation model, through text-conditioned pretraining and finetuning that aligns residues with their corresponding text. Prompteina outperforms evaluated baselines on 15 of 18 structural properties and generates proteins that more closely match descriptions of their intended functions. Prompteina outperforms state-of-the-art binder-design models on at least 13 of 20 targets by up to 14%, while supporting specifications beyond existing conditioning interfaces. With iterative in silico feedback, Prompteina’s gain over non-iterative generation is 57% larger than the next-best baseline on the hardest targets. Finally, we experimentally test Prompteina designs against GM2A, which has no reported de novo binders, and identify one binder whose binding signal increases with concentration and shows both association and dissociation. These results establish natural language as an effective and powerful conditioning interface for protein structure generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.