Rating-Scale Drift in Automated Paper Reviewers after Domain-Specific Reinforcement Learning
Abstract
Domain-specific post-training can improve an automated reviewer's agreement with its training labels, but does it preserve agreement with human ratings? We study this question by adapting a Qwen3.5-9B reviewer, initially fine-tuned on human reviews, with reinforcement learning against generated domain scores. In two source-matched Computer Science and Mathematics runs, active-domain Rating error decreases while human-referenced validation error increases by 43.3% and 29.3%, respectively. A separate offline evaluation compares these checkpoints and four additional, unmatched domain checkpoints with the same SFT baseline on jointly parsed human-rated papers. All six checkpoints have higher paired Rating error than SFT, accompanied by upward shifts or near-constant predictions. Some checkpoints also improve aggregate rule reward despite worse Rating agreement. The SFT baseline is itself concentrated, so these comparisons measure changes from a fixed, imperfect reviewer. We characterize these observations as rating-scale drift and report agreement, signed bias, and output concentration separately. The results document a separation between generated-domain progress and human-rating retention in the evaluated runs, without identifying its causal mechanism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.