acceptodds
Under review as a conference paper at ICLR 2027

MolDeTox: Evaluating Language Model’s Stepwise Fragment Editing for Molecular Detoxification

Abstract

Large language models (LLMs) are increasingly used for molecule design and optimization, but systematic evaluation of their capability for toxicity mitigation remains limited. We present MolDeTox, a benchmark for evaluating reference-based molecular detoxification across multiple fragment-level editing tasks. MolDeTox is built on ToxicityCliff, a collection of closely related molecular pairs that differ in toxicity for the same endpoint. These examples expose small structural changes associated with reduced toxicity while minimizing unrelated changes to the molecule. We convert these pairs into three QA tasks covering fragment identification, fragment substitution, and final molecule generation. This structure enables step-wise analysis of model behavior across the detoxification process. Our analysis further identifies common failure cases and highlights potential directions for improving performance. We also assess whether the benchmark provides a useful training signal. Fine-tuning a small language model on the training split leads to substantial gains on MolDeTox. The fine-tuned model also improves molecular repair success under an independent oracle-based evaluation setting. Together, these results show that MolDeTox serves not only as an evaluation benchmark, but also as a useful resource for developing and analyzing molecular detoxification systems. We release the evaluation code and dataset.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.