acceptodds
Under review as a conference paper at ICLR 2027

Not Just Saturation: Event-Structured Rewards for Open-Ended Post-Training

Abstract

During post-training for open-ended generation, judge scores for responses to the same prompt can concentrate within a narrow range as policy quality improves. We find that the issue is not the disappearance of reward differences, but whether the remaining differences within this narrow range reliably rank candidates by quality. When a limited judge struggles to distinguish closely matched responses, using these differences to estimate relative advantages can introduce noise or misdirect policy updates. Fine-grained rubric-based evaluation can reveal defects that holistic scores overlook, but directly aggregating the resulting diagnoses can introduce new problems. One underlying defect may be reported under multiple criteria, while several independent defects may be collapsed into a single criterion-level score. Poorly calibrated penalties can also overstate the impact of minor flaws and distort candidate rankings. We introduce CARMA, an event-structured reward framework that separates two questions: which diagnoses refer to the same defect, and how strongly each independent defect should affect reward. CARMA groups evidence-linked diagnoses that can be resolved by the same semantic repair into a single event and learns defect costs from same-prompt preferences. CARMA adjusts an independently obtained holistic score using event-level penalties, rather than assigning separate costs to multiple diagnoses of the same defect. In open-ended generation evaluations on Qwen3-8B and Qwen3.5-27B, CARMA-trained policies outperform the tested baselines in pairwise preference comparisons. It also yields more accurate near-tie rankings and advantage vectors that agree more closely with independently constructed references. Integrating event-structured signals into pairwise and feedback-driven training yields further gains over the corresponding methods without CARMA, supporting defect events as a reusable basis for open-ended post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.