DeskReject-Copilot: A Dataset, Approach, and LLM Benchmark for Submission Screening
Abstract
The number of submissions to journals and conferences has grown rapidly (e.g., ICLR: 2,213 in 2020 → 19,814 in 2026). This puts the peer review system under increasing strain. While prior work studies LLMs for writing reviews, their use for screening submissions for desk rejection has been identified as promising, but has neither been realized as a system nor evaluated empirically, in part because no labeled data existed. We introduce the first dataset of its kind: 1,104 ICLR desk rejections from 2020–2026, fully double-annotated with 20 rejection reasons. Building on it, we present DeskReject-Copilot, a detection pipeline that covers seven formal rejection reasons and provides a tailored configuration for each check. Instantiating it with ten multimodal LLMs released between May 2024 and August 2026 yields macro-averaged F1 scores between 0.75 and 0.99 in the balanced setting and between 0.71 and 0.97 in the imbalanced setting. The results reveal a marked asymmetry across desk-reject reasons: small open-weight models already achieve near-perfect detection for a subset of reasons, whereas for others, satisfactory performance is achieved only by a single closed-weight, high-cost model, with most current top-tier models falling short. Our findings suggest that DeskReject-Copilot can serve as an effective screening assistant for program chairs, albeit with a deliberate choice of model per desk-rejection reason. We release our dataset and code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.