Full Context Is Not Enough: Evidence-Grounded Meta-Review Generation
Abstract
As some major research venues begin to evaluate tightly bounded AI assistance while retaining strict limits on delegated reviewing, reliable evaluation of meta-review support is becoming increasingly consequential. Meta-reviewing must reconcile paper evidence, reviews, rebuttals, and unresolved disagreement into a defensible decision rationale; fluent automation can instead introduce unsupported paper claims or fictitious reviewer consensus. Existing benchmarks often expose only partial paper context and rely on reference-overlap metrics, leaving these failures difficult to measure. We introduce MetaBench, a 4,782-instance benchmark pairing full papers with review threads and human meta-reviews, and FACT, a pipeline that separates issue-conditioned planning from a return-to-source evidence gate. On a same-drafter 2×2 comparison over 1,520 cases, adding the evidence gate to Direct raises supported paper claims from 73.8% to 84.7%, whereas planning alone reaches 76.0%; Plan+Gate reaches 88.9% support and 80.6% major-issue coverage. On a 200-paper expert-gold slice, the same ordering holds, while claim extraction attains .79 precision, .82 recall, and .80 F1. A 120-case blinded study is descriptively favorable against the strongest gated baseline, although we make no confirmatory inferential claim. Across these evaluations, verification is associated with the largest support change, while planning adds most clearly to issue coverage and reductions in conflict errors when paired with the gate. Within the Direct stack, longer paper context produces smaller observed gains than explicit verification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.