H2L-Bench: A Benchmark for Agentic Planning Under Uncertainty in Drug Discovery
Abstract
AI agents have made rapid progress on tool use, multi-step planning and open-ended reasoning, yet the benchmarks that drive this progress rarely capture what matters in scientific decision making: acting under uncertainty, choosing between costly experiments, and committing to an answer with limited information. We introduce H2L-Bench, a lightweight decision-theoretic sandbox that simulates the hit-to-lead phase of small-molecule drug discovery as a budget-constrained sequential-decision task with a mechanistic ground truth. An agent chooses among eight assays of varying cost and information content, then nominates one compound from a library. Winning requires the compound's hidden traits to clear multiple thresholds simultaneously, with assay noise and budget scaled across three difficulty tiers. As a principled reference we implement a Bayesian value-of-information (VoI) agent that maintains an analytical Gaussian posterior over each compound's latent traits and picks the assay that most increases the best compound's probability of winning per dollar spent. Comparing VoI to four frontier LLMs (Claude Haiku 4.5, Sonnet 5, Opus 5, and OpenAI GPT-5.4) at 50 seeds per cell, we find that VoI has the highest win rate on every tier, but in paired same-seed tests its lead over the best LLM on each tier is not significant after correction for multiple comparisons. Beyond the headline win rates, the environment lets us analyse where LLMs and a Bayesian agent diverge action-by-action: on a subset of seeds, LLMs recover the correct compound by committing budget to a confirmatory profile of one compound, or to an expensive singular assay, that a depth-1 per-dollar VoI ranking keeps deferring. An approximate two-step lookahead does not close this gap on fresh seeds, so the shortfall reflects VoI's per-dollar scoring rule rather than its planning horizon. Test-time scaffolds do not help: on the hard tier, chain-of-thought and tree-of-thoughts leave Sonnet 5's win rate statistically unchanged, and explicit planning lowers it (32% to 23%).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.