SurveyAudit: Evidence-Grounded Detection of Reviewer Concerns in Literature Surveys
Abstract
The rapid advancement of large language models has made automated survey generation increasingly practical, yet evaluation of survey quality remains unreliable: existing methods rely on opaque LLM verdicts or heuristic metrics that lack empirical grounding and evidence transparency. We propose SurveyAudit, a verifiable framework focusing on the detection of reviewer concerns in literature surveys. We first introduce ConTax, an empirically grounded taxonomy of 11 recurring quality concerns derived from 519 real peer-review records on 168 survey papers. SurveyAudit then operationalizes these concerns through category-specific detection procedures and an evidence-tracing protocol: each detected concern must be supported by a verifiable source span, a bounded absence certificate, or a fully specified and reproducible computation path. Experiments on human-written, AI-generated, and perturbation-controlled surveys show that SurveyAudit substantially achieves higher precision and recall than LLM- and tool-augmented judges on reviewer-relevant concern detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.