SciXbench exposes the fragility of LLMs in long-horizon scientific problem solving
Abstract
A plausible scientific answer can conceal a flawed research process. We introduce SciXbench, a benchmark of 83 expert-curated, long-horizon research tasks across seven scientific subdomains, centered on chemistry. Beyond task construction, we make two methodological contributions to scientific evaluation: (1) Dynamic Scientific Rubrics (DSRs). We design a literature-grounded method that builds expert-reviewed subdomain frameworks and uses a multi-agent workflow to derive question-specific scoring criteria, scales, and weights while preserving criterion provenance. (2) Scientific Trace. We develop a non-intrusive approach to recording and evaluating scientific trajectories. It captures native runtime events, tool outputs, and generated artifacts without prescribing agent workflows, and links decisions, observations, and final results through their underlying evidence. DSR-based answer scores and trajectory scores are computed independently and analyzed jointly to diagnose execution errors, misinterpretation, and failures to carry evidence across research stages. On 15 FrontierScience chemistry problems, the selected DSR configuration achieves 76.03% pairwise agreement with expert-rubric judgments. Trace-derived supervision further improves strict top-5 retrosynthesis accuracy by 4.51% on a deduplicated ORDerly subset. SciXbench provides both challenging scientific tasks and a methodology for evaluating the outcomes and processes of scientific problem solving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.