Calibrating Rubric Partition Sensitivity in LLM Judges
Abstract
Equivalent rubric content need not produce equivalent LLM-judge scores when its criteria are grouped differently. Raw differences confound this interface effect with repeat variability, score discreteness, and the number of partitions tested. We introduce an estimator-matched conditional randomization calibration and apply it to 780 items in seven dataset cohorts, five corpus families, and 96,120 successful calls across DeepSeek v4 Pro, Qwen3.8 27B, and Selene-1-Mini. Direct grouped scoring has calibrated drift ratios CDR = 1.259, 1.239, and 2.355 (95% intervals reported in that model order). Item-level excess remains positive after 10% within-cohort trimming for all three models, and no single item changes the corresponding pooled direct-excess estimate by more than 10%. Post-hoc mechanism diagnostics show signed residual centralization across the score range, additional multi-atom contextual residual on common support, and conditional associations with block size and atom disagreement. A partition-independent atom-fixed control has no comparable positive score-unit excess for any model. The holistic single-block contrast is model-dependent, underscoring that the result is a measurement property of the scoring interface rather than a correction coefficient that transfers unchanged across evaluators.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.