acceptodds
Under review as a conference paper at ICLR 2027

Molecular Structure Elucidation with Frontier Models: A Benchmark on Literature-Reported IR, H, and C NMR Peak Lists

Abstract

Frontier models are often presented as near-solved structure elucidators on curated spectra. We ask an operational question: given the molecular formula and literature-reported IR, H, and C NMR peak lists, printed wavenumbers and shift tables, not digitised spectra images, can an out-of-the-box frontier model in an agentic workflow recover the correct constitution, and which stage fails when it does not? We release IRSpectra-Bench as a frozen roster of problems, an RDKit InChIKey-14 contract, and the prediction deposits a mechanical scorer can replay. The object is a factorised diagnosis of one closed-book harness. One trained NMR-only baseline is scored on the same roster: NMRTrans, formula and printed H/C lists, IR unused (79/500 top-1; Appendix A.13). It is not a full multi-modal peer of the IR+H+C task. On one closed-book frontier harness, generation top-1 is 45.4% (227/500) and recall (the fraction of problems whose true structure enters the proposed pool) is 49.8% (249/500). Self-ranking places the true structure first in 91.2% of the cases (). The generation wall is 227 exact top-1 / 22 in-pool not top-1 / 251 never proposed. The roster is near-50/50 simple/complex; reweighted to the eligible corpus mix, top-1 on a validate-clean subset is 26.5%. Forward verification of the same pools is a secondary drop, not a second wall: top-1 falls from 227/500 to 204/500 (; diagnostic ; Appendix A.8). On a 60-compound arm only, formula-only recovery is 3/60, accuracy is flat in source-paper year, and recall stays below precision across four vendor families.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.