acceptodds
Under review as a conference paper at ICLR 2027

SafeProtein++: Cross-Architecture Red-Teaming and Unlearning of Protein-Generative Models

Abstract

Proteins execute nearly every cellular function, and neural networks now design them from a multi-modal context—a functional label, homology sequences or a backbone structure—denoising sequence and structure jointly or decoding one amino acid at a time. These models deliver functional enzymes, binders and antigens, but their susceptibility to jailbreaking is underexplored, raising the concern of harmful outputs such as toxins and viral proteins. We introduce SafeProtein++, the first framework to evaluate the jailbreak vulnerabilities of protein generative models across all three main architecture families and the safeguards proposed to close them. It provides (1) a cluster-disjoint benchmark of 854 proteins supplying identical cells to all three generator families, judged jointly on sequence and folded structure; (2) a pathogenic-guided beam search that steers chunked decoding toward high-hazard completions; and (3) a reinforcement-unlearning recipe that removes hazardous capability from the weights at affordable cost in benign generation. Empirically, masking conserved positions is the strongest attack for every family, and unlearning cuts hazardous attack success to a few percent on every generator, at a measurable cost in structural fidelity. The guided beam lifts attack success on every mask condition, yet against an unlearned model it collapses to a fraction of the base rate: search-time guidance re-ranks the victim’s own distribution rather than restoring removed hazard. A case study on amylin reveals a hazard invisible to every guardrail: completing its masked amyloid core raises predicted aggregation propensity above the native hormone. Known-target reconstruction is therefore insufficient for safety certification; despite unlearning reduces covered hazards, while open-set coverage remains unresolved. Code and data will be released publicly upon paper acceptance, and an anonymized version is available at https://anonymous.4open.science/r/ICLR-2027-SAFEPROTEIN-PP/README.md.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.