Student Geng’s WorkBuddy: Evidence-Grounded Verification of Scientific Concerns
Abstract
Post publication peer review supports scientific self-correction, but systematic, reproducible verification of public scientific concerns still relies heavily on manual effort. PubPeer is an open platform where researchers question published papers, share supporting material, and discuss the reliability of the work. It has consequently become a large collection of research concerns. Recently, Hongwei Geng, also known as Student Geng, used public PubPeer comments and related evidence to examine data anomalies in papers from leading journals, such as *Nature*. The resulting cases led to investigations and sanctions by organizations such as the National Natural Science Foundation of China. Such cases show that PubPeer comments provide scientifically and practically valuable leads for research scrutiny, but remain unverified and require substantial manual validation. To address this problem, we propose Student Geng's WorkBuddy-27B, an evidence verification assistant for PubPeer comments. For each comment and corresponding paper, WorkBuddy first reformulates the scientific concern as a neutral statement, removing subjective or emotional language while preserving verifiable factual claims. It then maps the concern to traceable evidence in paper text, figures, methods, supplementary materials, and source data. WorkBuddy assesses how strongly this evidence supports the concern, together with evidence quality, potential impact, and evidence location. For training and evaluation, we construct PubPeerBench-1000 with 1,000 cases annotated by experts. We train Student Geng’s WorkBuddy through supervised fine-tuning and compare it on an independent test set with leading large language models, including GPT-6 Astra. Against untuned Qwen3.6-27B, our model's verdict exact agreement rises from 38.50% to 48.00%, Macro-F1 from 28.13% to 37.00%, EQ MAE falls from 1.50 to 0.89, and impact exact agreement rises from 23.50% to 39.50%, showing the effectiveness of task specific fine tuning. *PubPeerBench-1000 and code will be publicly released.*
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.