acceptodds
Under review as a conference paper at ICLR 2027

Can Anyone Make Poisons? Systematic Red-Teaming of Molecule Language Models

Abstract

Molecule language models offer an accessible interface for molecule design, but also raise potential dual-use concerns. Despite extensive research on red-teaming large language models, the safety risks of these models remain largely unexplored. In this work, we present the first systematic red-teaming study of molecule language models. We curate toxicity data spanning eight endpoints and develop specialized evaluators for systematic assessment. Our analysis further reveals that toxic molecules exhibit greater similarity to known toxic examples than non-toxic molecules. Building on this structural prior, we develop a decoding-time red-teaming strategy that learns a surrogate scorer to estimate the toxic-neighbor potential of partial SMILES and steer generation toward toxicity-associated regions. Extensive experiments across eight models show that our method substantially increases toxicity rates, yielding an average absolute improvement of up to 41.4%. Qualitative analyses further identify generated compounds with established real-world toxicological hazards. These findings expose an overlooked attack surface in molecule language models and underscore the need for stronger safeguards, monitoring, and traceability mechanisms for frontier systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.