DRYLAB-Bench: Benchmarking the Biosecurity Risk of Agentic LLM Mutation Design
Abstract
As large language models (LLMs) acquire advanced biological design capabilities, they offer substantial benefits for research but simultaneously introduce severe dual-use risks, such as the engineering of enhanced pathogens or drug-resistant bacteria. While international biosecurity frameworks recognize these threats, the actual risk posed by LLMs in executing dangerous biological modifications remains unquantified, particularly when models are augmented with specialized biological design tools (BDTs). Existing evaluations primarily assess the possession of hazardous knowledge rather than its practical application, overlook the risk amplification from BDTs, and rely on validation methods that are either unreliable, prohibitively expensive, or inherently hazardous. To address these gaps, we introduce DRYLAB-Bench (Design Risk Yielded by LLM Agent in Biodesign), the first comprehensive benchmark designed to evaluate the high-risk biological mutation design capabilities of LLM agents. Moving beyond basic question-answering, DRYLAB-Bench tasks agents with 30 genuinely hazardous mutation-design scenarios distilled from biosecurity guidelines. It evaluates models in a multi-round, tool-augmented environment where LLMs orchestrate BDTs to iteratively refine candidates. Crucially, it scores submitted designs directly against 187,695 empirical wet-lab measurements from ProteinGym and ViroGym through our proposed Biosecurity Risk Score (BRS), which quantifies the risk-associated effects of submitted mutations while enabling scalable and safe ground-truth verification. Evaluations of 15 frontier models reveal substantial baseline threats: without tools, 67.6% of one-shot designs exceed the average measured variant, and 21.4% reach the highest-risk decile. Equipping models with BDTs significantly amplifies this risk, raising the pooled risk by ∆BRS = +0.116 and the top-decile share to 23.5%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.