ToxRedTeam: Red-Teaming Protein Toxicity Detectors with Structure-Aware Adversarial Attacks
Abstract
Toxic protein detection plays a critical role in screening DNA synthesis orders. Existing methods generally fall into two categories: homology-based detection and machine learning (ML)-based detection. Prior work has shown that homology-based approaches can be bypassed through strategically designed mutations to toxic proteins. In contrast, ML-based detectors learn sequence representations rather than relying on explicit matching, but their robustness to adversarially modified toxic protein variants remains largely unexplored. To address this gap, we propose ToxRedTeam, a systematic red-teaming framework for auditing protein toxicity detectors across both categories. Given query access to a detector’s prediction API, ToxRedTeam iteratively generates variants of a toxic protein using a masked protein language model, with the objective of inducing misclassification as non-toxic while preserving the protein’s toxic functionality. We empirically demonstrate the effectiveness of ToxRedTeam across five detectors, where it consistently outperforms baseline attack methods. Our code will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.