BRIDGE-bench: Benchmarking Real-world Issues in Drug Generation and Evaluation
Abstract
Large language models (LLMs) and tool-using agents are increasingly proposed as assistants for drug discovery, but their evaluation remains fragmented. Existing benchmarks often emphasize isolated capabilities, such as molecular property prediction and chemical question answering, while practical therapeutic research requires coordinated decisions across biology, chemistry, pharmacology, modeling, manufacturability, and translation. We introduce BRIDGE-Bench (Benchmarking Real-world Issues in Drug Generation and Evaluation), a pipeline-level benchmark built from real-world drug discovery and development questions. BRIDGE-Bench covers target biology, hit and lead discovery, binding pocket analysis, absorption, distribution, metabolism, and excretion (ADME), quantitative systems pharmacology (QSP), retrosynthesis, antibody-drug conjugate (ADC) design, drug delivery, chemistry, manufacturing, and controls (CMC), and wet-lab planning. Each instance is organized around a concrete decision scenario with reference answers and scoring criteria. We evaluated 13 harness-model configurations within the Claude Code and Codex agent frameworks. Final average scores ranged from 19.3 to 36.0, with the strongest aggregate result obtained by Claude Code with claude-opus-5. The benchmark supports a modular and trajectory-level evaluation of answer correctness, evidence grounding, tool use, uncertainty awareness, cross-stage trade-off reasoning, and error modes that matter for practical therapeutic decisions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.