PaperRecon: Characterizing Scientific Writing Capabilities of AI Agents through Paper Reconstruction
Abstract
As coding agents become capable of generating entire research papers, assessing their actual writing quality and risks has become critical. Existing evaluations primarily rely on overall review scores, which assess acceptance-worthiness but provide limited insight into how individual agents write. We introduce (PaperRecon), a novel framework for characterizing scientific writing capabilities from multiple perspectives. PaperRecon formulates scientific writing as a reconstruction task. Given a structured overview and supporting resources derived from an existing paper, an agent reconstructs the full paper, which is then directly compared with the original. We evaluate two core aspects of scientific writing. measures how well the reconstructed paper captures the intended scientific content using section-specific rubrics, while analysis identifies claims that contradict the source paper and remain unsupported after checking the provided resources. We further conduct four diagnostic analyses to characterize the effects of the harness and backbone model, input information, citation behavior, and verbosity. To support this evaluation, we introduce , comprising 72 recent papers from top-tier venues in 2025 and 2026, filtered to reduce contamination from prior model knowledge. Our evaluation reveals clear differences across model families. Claude-based agents achieve higher content coverage than GPT-based agents but produce substantially more hallucinations. These results show that scientific writing capability cannot be adequately characterized by a single overall score and demonstrate how PaperRecon provides fine-grained insights into the strengths and weaknesses of current coding agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.