OrganicBench: Can Large Language Models Reason Like Chemists?
Abstract
Large language models have surpassed human experts in coding and mathematical theorem proving. However, they still struggle to tackle natural science problems in the real world, thereby limiting their potential for autonomous scientific discovery. We find that organic chemistry is a rich and challenging testbed grounded in real-world chemistry for evaluating AI chemists. To this end, we introduce OrganicBench, a multimodal benchmark comprising over 1,200 challenging organic chemistry problems grounded in published research literature. Given problem descriptions in images, SMILES strings, or molecular graphs, models are tasked with predicting reaction products, inferring key intermediates, and reasoning about reaction mechanisms. These tasks require models to understand realistic reaction conditions, apply correct knowledge, and perform complex reasoning in multistep synthetic routes, going beyond factual question answering based on memorization. Our evaluation demonstrates that frontier proprietary and open-source models fail to achieve competitive performance on these problems, and they exhibit consistent failure modes. Process-level analyses further reveal failures across multiple stages of problem-solving, exposing the gap between broad chemical knowledge and reliable complex reasoning. Overall, OrganicBench serves as an expert-level benchmark for measuring progress toward capable AI chemists.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.