RubricOpt: Recursive Refinement of Semantic Rubrics for Preference Optimization
Abstract
Pairwise preference labels specify which response is preferred but leave the behavioral basis of that preference implicit. We introduce RubricOpt, a framework that makes a semantic rubric an explicit, editable specification for preference optimization. RubricOpt measures semantic demand—how behaviors separate task success from failure—and semantic supply—how often preference pairs contrast those behaviors—in a shared eight-factor space. Semantic Rubric Steering uses this diagnosis to condition training on a short natural-language rubric while keeping the response pairs, labels, and preference objective fixed. The rubric is gradually withdrawn before training ends, and all evaluation is rubric-free. A gated refinement loop revises clauses when recurring failures reveal requirements missing from the current wording. Our analysis gives sufficient conditions for criterion-free behavioral improvement and for retaining success-probability gains after withdrawal. Experiments use two preference corpora and three models from two families, each compared with a matched ordinary-DPO control. On Tulu3, steering improves BBH by – percentage points and the seven-benchmark average by – points across Qwen3-4B, Qwen3-8B, and Mistral-7B. On HelpSteer3, four rubric compositions improve Qwen3-4B Arena-Hard win rates by – points, without a material increase in response length. Matched HelpSteer3 continuation on Qwen3-4B adds BBH points under the unchanged rubric, while the tested wording edit yields no detectable additional BBH gain, separating the benefit of continued training from the effect of refinement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.