acceptodds
Under review as a conference paper at ICLR 2027

SuperResearch Bench: Stress-Testing Long-Horizon Research with Evidence Graphs and Rubric-Guided Audits

Abstract

While Large Language Models (LLMs) have demonstrated proficiency in Deep Research or Wide Search, their capacity to solve highly complex questions requiring long-horizon planning, massive evidence gathering, and synthesis across heterogeneous sources remains largely unexplored. We introduce SuperResearch Bench, a benchmark for Super Research: complex autonomous research tasks that integrate (i) structured decomposition, (ii) super wide retrieval for diverse perspectives, and (iii) super deep investigation through iterative queries. To evaluate this capability, we curated 300 expert-written questions across diverse domains, each requiring up to 100+ retrieval steps and 1,000+ web pages to reconcile conflicting evidence. SuperResearch Bench provides source-linked evidence graphs, reference reports with fine-grained citations, and intermediate artifacts to support traceable reasoning. The benchmark artifacts undergo automated validation and expert review. Our graph-anchored, rubric-guided auditing protocol evaluates generated reports along four primary dimensions: Coverage, Logical Consistency, Report Utility, and Objectivity. Citation Health is reported separately. While such questions may be infrequent in standard applications, SuperResearch Bench serves as a ceiling evaluation and stress test for long-horizon research. A comparison of fourteen systems and protocols reveals substantial variation in evidence recovery, reasoning consistency, and report usefulness, enabling fine-grained diagnosis of current research systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.