Translating In Silico Toxicology: Benchmarking Large Language Models for Interpreting Machine Learning-Driven Drug Discovery
Abstract
Clinical trials of potential new medicines are exceptionally costly and often end in failure, with approximately 30% of these failures attributed to drug toxicity. In silico drug toxicity prediction has been an ever-growing field in recent years. However, computational toxicity models often lack interpretability for drug design specialists. In this paper, we construct a framework using Large Language Models (LLMs) to interpret machine learning toxicity predictions and feature importances to clearly communicate the outputs to drug design specialists. Using a PubChem derived dataset, we test four LLMs across multiple structured prompts for varying tasks. These outputs are evaluated using a rigorous three-tier framework; deterministic evaluation, error analysis, and LLM-as-a-judge. The evaluation of the framework identified Gemma3 as the best performing model according to deterministic metrics, while LLM-as-a-judge credited Deepseek-r1 with superior reasoning depth and structural verifiability. In addition, performance was seen to be robust to random variation. Crucially the error analysis identified zero occurrences of feature hallucinations or magnitude distortions, thus establishing the reliability of LLM-generated interpretations for drug discovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.