acceptodds
Under review as a conference paper at ICLR 2027

ChemDialog: Benchmarking Chemical Reaction Understanding in LLMs through Dialogue

Abstract

Despite the promise of large language models (LLMs) for chemical reaction understanding, current benchmarks rely heavily on question-answering (QA) for isolated tasks, such as molecule property prediction and stoichiometry. However, real-world chemistry applications require LLMs to engage in interactive dialogue and integrate information across related tasks. To address this gap, we introduce ChemDialog, a reaction-grounded dialogue benchmark for evaluating LLMs’ ability to integrate chemical information across dialogue turns and related tasks. Specifically, we construct ChemDialog by organizing chemical reaction subtasks into coherent dialogue scenarios with LLM assistance, followed by extensive human review for quality control. We further develop a multidimensional evaluation framework that separately assesses task-specific chemical competence and dialogue-level interaction logic. Extensive experimental results indicate that mainstream LLMs still struggle with logical consistency and dynamic information integration in multi-turn chemical dialogues, and providing multi-turn, related task dialogue history significantly enhances model performance in synthesis tasks. We hope that ChemDialog and our findings encourage the research community to further advance the untapped potential of LLMs toward complex, real-world chemical interactions. The dataset and code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.