acceptodds
Under review as a conference paper at ICLR 2027

OneResearchBench: Benchmarking Research Agents for Research Synthesis from Real-World Source Materials

Abstract

Progress toward general-purpose AI agents requires capabilities beyond isolated question answering. Agents must operate reliably across diverse domains, handle complex user-provided materials, and complete long-horizon knowledge tasks that involve gathering evidence, integrating information, and producing trustworthy outputs. Existing research-agent benchmarks, however, largely start from research questions or open information needs, and give limited attention to research that begins from real-world source materials. We define source-grounded research synthesis, a task in which an agent must first understand provided materials, then conduct external research, integrate the resulting evidence, and produce a grounded long-form report. To evaluate this capability, we introduce OneResearchBench, containing 101 human-constructed tasks across six domains and 3,671 expert-written, task-specific rubric rules. The rubrics evaluate source understanding, external research, evidence grounding, technical accuracy, citation traceability, and structure quality. Our criterion-level evaluator achieves a Pearson correlation of 0.909 with expert scores on the human calibration subset. We evaluate eight model–harness configurations. Under direct report generation, no system scores above 44.75/100, with external research and citation traceability emerging as the largest capability gaps. A structured skill-driven research workflow improves Claude Opus 4.8 with Claude Code from 44.75 to 53.07 (+8.32), and GPT-5.5 with Codex from 43.03 to 47.90 (+4.87). These results show that source-grounded research synthesis remains challenging for current research agents. OneResearchBench provides a benchmark for measuring progress on this task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.