Rubrics Always Miss: Robust Preference Optimization under Incomplete Criteria
Abstract
Rubrics provide fine-grained and interpretable supervision for aligning large language models by evaluating responses against explicit criteria, such as factual correctness, relevance, and instruction following. Recent methods generate rubrics that specify what makes a good response to a given instruction, then use rubric scores as rewards or to construct preferences for training. However, it is difficult to ensure that these rubrics cover all relevant criteria, so their scores may not fully reflect response quality. To address this issue, we propose a robust preference optimization framework that accounts for errors in preferences caused by incomplete rubrics. Since the missing criteria are unobserved, we assess rubric completeness indirectly by examining whether the observed rubrics capture differences between responses. For each instruction, we perturb the policy to generate response variations and measure how consistently larger response changes are accompanied by larger changes in rubric scores. This agreement serves as an instance-specific proxy for rubric completeness. We then use the estimated completeness to construct a feasible set around the preference probabilities inferred from the observed rubric scores, allowing larger deviations for samples with lower estimated completeness, and therefore a larger feasible set. The policy minimizes the worst-case training loss over this set, thereby allocating greater protection against preference errors to samples with less complete rubrics. Our framework can be integrated with any margin-aware preference optimization methods. Experiments across multiple datasets demonstrate consistent improvements over existing methods. Code is on anonymous GitHub.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.