acceptodds
Under review as a conference paper at ICLR 2027

BRIDGE-bench: Benchmarking Real-world Issues in Drug Generation and Evaluation

Abstract

Large language models (LLMs) and tool-using agents are increasingly proposed as assistants for drug discovery, but their evaluation remains fragmented. Existing benchmarks often emphasize isolated capabilities, such as molecular property prediction and chemical question answering, while practical therapeutic research requires coordinated decisions across biology, chemistry, pharmacology, modeling, manufacturability, and translation. We introduce BRIDGE-Bench (Benchmarking Real-world Issues in Drug Generation and Evaluation), a pipeline-level benchmark built from real-world drug discovery and development questions. BRIDGE-Bench covers target biology, hit and lead discovery, binding pocket analysis, absorption, distribution, metabolism, and excretion (ADME), quantitative systems pharmacology (QSP), retrosynthesis, antibody-drug conjugate (ADC) design, drug delivery, chemistry, manufacturing, and controls (CMC), and wet-lab planning. Each instance is organized around a concrete decision scenario with reference answers and scoring criteria. We evaluated 13 harness-model configurations within the Claude Code and Codex agent frameworks. Final average scores ranged from 19.3 to 36.0, with the strongest aggregate result obtained by Claude Code with claude-opus-5. The benchmark supports a modular and trajectory-level evaluation of answer correctness, evidence grounding, tool use, uncertainty awareness, cross-stage trade-off reasoning, and error modes that matter for practical therapeutic decisions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.