acceptodds
Under review as a conference paper at ICLR 2027

MolP2SBench: Do LLMs Really Reason from Molecular Properties to Structure?

Abstract

Large language models (LLMs) are increasingly used in molecular discovery and structure elucidation, where decisions depend on physical observations. Existing evaluations primarily focus on end-to-end molecular reconstruction, where correct predictions do not necessarily indicate that these results are grounded in the supplied properties. To address this, we introduce MolP2SBench, a benchmark for evaluating molecular property-to-structure (P2S) reasoning in LLMs. Each task presents a masked molecular scaffold with given properties, requiring the model to select the missing structure from plausible candidates. To assess whether predictions genuinely rely on the provided properties, we test how models respond when properties are shuffled or replaced, thereby evaluating their reliance on correct property-structure pairing and their responsiveness to counterfactual property changes. We also establish a channel-wise leaderboard using Property-Grounded Recovery (PGR), which gives equal weight to source and counterfactual recovery. In our evaluation of seven models, we find that LLMs can leverage physical properties for molecular structure prediction, but their performance varies substantially across property types. More importantly, higher accuracy with a given property does not necessarily imply greater responsiveness to that property. For example, a model may benefit more from H NMR than from C NMR, yet be less likely to change its prediction when the H data is replaced. Interestingly, LLM performance does not necessarily reflect the relative difficulty of these properties for human chemists. For instance, although UV–Vis spectra provide less direct structural information than NMR spectra, Astra achieves higher accuracy with UV–Vis data than with NMR peaks. In summary, MolP2SBench provides a framework for studying how LLMs extract and use information from different physical properties, revealing performance patterns that may differ from expectations based on conventional chemical intuition and offering a foundation for developing more reliably property-grounded models for molecular and quantum-chemical inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.