acceptodds
Under review as a conference paper at ICLR 2027

Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation

Abstract

LLM judges increasingly evaluate responses against fine-grained rubric checklists. Evaluating all rubrics in a single pass is efficient, but we find it introduces *rubric interference*: the verdict on one rubric shifts depending on which other rubrics are co-evaluated. In a preliminary study, only one-third of samples receive fully consistent verdicts across rubric sets of varying composition. We formalize this problem and develop a measurement framework that probes interference through three controlled operations: rubric set expansion, reordering, and noise injection. To mitigate interference without external supervision, we propose **Self-Anchored Rubric Alignment (SARA)**. SARA uses a model's own single-rubric judgments as stable anchors and aligns multi-rubric outputs to these anchors through on-policy self-distillation. We validate SARA on three datasets (HealthBench, FLASK, ResearchQA) across Qwen3 and Llama-3.1 model families. SARA improves rubric-level consistency by up to 16% and triples sample-level exact match, while preserving agreement with both the base model and GPT-4.1. We further show that judge consistency directly affects downstream reinforcement learning: a SARA-aligned reward model reduces inter-run reward variance by 5x under rubric-order perturbation, confirming that consistent rubric evaluation is a practical requirement for stable policy optimization. Our code is available at https://anonymous.4open.science/r/SARA-22E2/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.