acceptodds
Under review as a conference paper at ICLR 2027

Rubric-to-Segment: Criterion-grounded Semantic Credit Assignment for Language Model Alignment

Abstract

Rubric-based reinforcement learning translates open-ended alignment objectives into structured, interpretable supervision through explicit evaluation criteria. However, standard rubric-based group relative policy optimization (GRPO) aggregates criterion scores into a single reward and broadcasts the resulting advantage across the response, obscuring criterion-specific performance and its correspondence to response content. Our analysis of safety and medical rubrics and model responses finds that most criteria concern specific semantic content, with evidence often localized to different parts of a response. Building on this correspondence, we introduce Rubric-to-Segment, which turns structured rubric judgments into contentspecific training feedback by linking criterion-wise comparison across responses with evidence-based allocation within each response. Using scores and evidence from a single judge assessment, it computes group-relative advantages separately for each criterion and allocates them to semantic segments as local corrections to the response-level advantage. Across two backbones in safety and medical alignment, Rubric-to-Segment improves task performance over response-level rubric GRPO-based and reinforcement learning–distillation baselines while maintaining general capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.