acceptodds
Under review as a conference paper at ICLR 2027

RITE: Investigating and Mitigating Rubric Drift in LLM-Based Grading

Abstract

Large language models are increasingly used to grade open-ended student responses, yet construct-irrelevant variation can change their judgments. We call this failure rubric drift: grading changes despite fixed questions, rubrics, and rubric-relevant evidence. We introduce EduR²Bench, with 1,060 verified rubric-equivalence groups and 3,297 variants across four STEM subjects, and Criterion Flip Rate (CFR), a label-free drift diagnostic. Rubric drift is widespread across open-weight, frontier, and specialized graders and is distinct from grading accuracy. Our central finding is supervision-induced invariance (SII): correctness supervision on original responses reduces CFR on unsupervised equivalent variants from 26.3% to 11.7%, accompanied by contraction of upper-middle-layer representation gaps. However, SII is limited by supervision coverage, label reliability, and transformation-family transfer. These boundaries motivate Rubric-Invariant Training via Equivalence (RITE), which combines correctness grounding with equivalence supervision. Using labels only on original responses, RITE reduces CFR to 5.2% while preserving grading quality and retains substantial benefit on a held-out transformation family and unseen human-written rewrites. Together, our results characterize when rubric invariance emerges from supervision and how to strengthen it explicitly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.