acceptodds
Under review as a conference paper at ICLR 2027

SciXbench exposes the fragility of LLMs in long-horizon scientific problem solving

Abstract

A plausible scientific answer can conceal a flawed research process. We introduce SciXbench, a benchmark of 83 expert-curated, long-horizon research tasks across seven scientific subdomains, centered on chemistry. Beyond task construction, we make two methodological contributions to scientific evaluation: (1) Dynamic Scientific Rubrics (DSRs). We design a literature-grounded method that builds expert-reviewed subdomain frameworks and uses a multi-agent workflow to derive question-specific scoring criteria, scales, and weights while preserving criterion provenance. (2) Scientific Trace. We develop a non-intrusive approach to recording and evaluating scientific trajectories. It captures native runtime events, tool outputs, and generated artifacts without prescribing agent workflows, and links decisions, observations, and final results through their underlying evidence. DSR-based answer scores and trajectory scores are computed independently and analyzed jointly to diagnose execution errors, misinterpretation, and failures to carry evidence across research stages. On 15 FrontierScience chemistry problems, the selected DSR configuration achieves 76.03% pairwise agreement with expert-rubric judgments. Trace-derived supervision further improves strict top-5 retrosynthesis accuracy by 4.51% on a deduplicated ORDerly subset. SciXbench provides both challenging scientific tasks and a methodology for evaluating the outcomes and processes of scientific problem solving.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.