acceptodds
Under review as a conference paper at ICLR 2027

MolViBench: Evaluating LLMs on Molecular Vibe Coding

Abstract

*Molecular Vibe Coding*, a paradigm where chemists interact with LLMs to generate executable programs for molecular tasks, has emerged as a flexible alternative to chemical agents with predefined tools, enabling chemists to express arbitrarily complex, customized workflows. Unlike general coding tasks, molecular coding imposes a distinctive challenge that LLMs should jointly equip programming, molecular understanding, and domain-specific reasoning capabilities. However, existing benchmarks remain disconnected. General code generation benchmarks such as HumanEval and SWE-bench require no chemistry knowledge, while chemistry-focused benchmarks such as S-Bench and ChemCoTBench evaluate knowledge recall or property prediction rather than executable code generation. To bridge this gap, we introduce **MolViBench**, the first framework and benchmark tailored for Molecular Vibe Coding. MolViBench comprises 358 curated tasks across five cognitive levels, ranging from single-API recall to end-to-end virtual screening pipeline design, spanning 12 real-world drug discovery workflows. To rigorously assess generated code, we further propose a multi-layered evaluation framework that combines type-aware output comparison and AST-based API-semantic fallback analysis, which jointly measures executability and chemical correctness. Evaluating 9 frontier LLMs under direct generation, incremental repair, and agent collaboration reveals a sharp gap: models succeed on API recall but fail on end-to-end molecular workflow synthesis, with all models below 10% Pass@1 on Level 5. These results position MolViBench as a diagnostic testbed for measuring executable chemical reasoning in AI-assisted molecular discovery.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.