PaperEvaluation: Benchmarking Agent-Generated Scientific Papers with Execution Evidence
Abstract
Large language models are increasingly capable of performing complex, long-horizon research tasks. However, existing scientific benchmarks provide limited coverage of end-to-end workflows from data analysis and experimental validation to complete manuscript generation, and lack targeted evaluation that connects research execution to the final manuscript. We introduce PaperEvaluation, a benchmark for evaluating agent-generated scientific papers using execution evidence. The benchmark comprises 80 research task packages derived from published papers across 13 research domains. Each task provides a research question, the necessary data, and a controlled execution environment while withholding the source paper and its conclusions. Agents must independently conduct the research and produce a complete manuscript. To evaluate this long-horizon research process, we propose R3Eval, a framework with three dimensions: Research Completeness, Reporting Rigor, and Reasoning Alignment. Research Completeness assesses whether the required scientific work is supported by execution evidence. Reporting Rigor examines whether the manuscript accurately reports its methods and findings. Reasoning Alignment compares the generated and source papers for functional alignment of their scientific arguments. Finally, we evaluate 14 agent systems based on PaperEvaluation, revealing distinct performance profiles across these dimensions through an evaluation of long-horizon research that connects execution to the final manuscript.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.