Extending Controllability from Composition and Topology to Domain Identity in Language-Model Protein Generation
Abstract
Protein language models can generate plausible sequences, but reliable control over multiple user-specified properties remains limited. We present Aria, which tunes Llama-3.1-8B-Instruct on 350k Swiss-Prot proteins described by a random subset of thirteen physicochemical, developability, structural, and functional descriptors written into the prompt, and which measures adherence by recomputing every descriptor from the generated sequence using the same code that wrote the prompt. On cluster-separated evaluation proteins, Aria improves adherence across directly measurable properties, including length, hydrophobicity, cysteine count, transmembrane topology, and signal peptides. Prompt-mismatch controls show that these gains depend on the requested specification. Exact Pfam-domain recovery remains difficult and increases with family exposure. We next test whether natural-sequence-anchored direct preference optimization can improve weaker structural and functional axes. One optimization stage raises mean ESMFold pLDDT by 8.2 points in PF00416, ribosomal protein uS13, and increases exact requested-domain recovery by 23.6-46.0 percentage points across three additional optimized families, with corresponding increases in predicted structural similarity. However, applying the resulting family-specific adapters to six families not used for preference optimization does not improve target-domain recovery over the supervised model. These results show that a general instruction-tuned model can acquire broad protein-property control without protein-specific pretraining, while natural-sequence-anchored preference optimization provides substantial but family-local gains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.