FREA: A Multi-Source Expert Benchmark for Reaction Feasibility Verification
Abstract
Reaction feasibility verification is essential for synthesis planning, yet progress is constrained by limited evaluation resources and scarce negative data. We introduce FREA, an expert-annotated benchmark of reactions spanning retrosynthesis proposals, zero-yield experimental records, perturbations by large language models (LLMs), and candidates from five generation methods. This multi-source design enables a broad assessment of verifier strengths and failure modes. Each reaction is manually annotated by expert chemists under carefully specified criteria. We evaluate frontier LLMs, forward models, and learned classifiers on FREA. Frontier LLMs show strong performance, with the best attaining the highest mean balanced accuracy across sources and closely matching expert rankings of retrosynthesis models. Forward models achieve the best performance on retrosynthesis proposals, with likelihood scoring providing an efficient option for large-scale screening. However, no model performs best across all four sources, and each exhibits distinct failure patterns. We also release a corpus of over 14 million recorded reactions and generated negative candidates. Exploratory training experiments suggest some transfer from generated negatives to other reaction sources, and generating negative data that reliably improve generalization remains an open question. Together, the benchmark and corpus provide a shared platform for diagnosing verifiers and studying how to improve them.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.