acceptodds
Under review as a conference paper at ICLR 2027

A Preference Is a Spectrum, Not a Scalar: PRISM for Structured Ordinal Reward Modeling

Abstract

Pairwise preferences record which response wins, but discard information an annotator often knows: which aspects separate the responses, how large the gap is, and why the rejected response fails. We study whether this discarded structure, rather than model capacity, limits reward modeling. We re-audit the Skywork-80K preference prompts into five aspect ratings, an ordinal preference strength, and a 17-way error taxonomy, yielding 66,605 valid pairs. We then train PRISM (Preference Refraction Into a Structured Mixture), a lightweight head on a frozen public reward model. PRISM uses sample-conditioned reward atoms, per-aspect ordinal distributions, and a coupled pairwise likelihood with explicit tie mass. A separate prompt-gated verifier handles programmatically checkable constraints without changing rankings when the gate is inactive. On RewardBench 2, the trained head alone reaches 0.855 with Llama-3.1-8B and 0.814 with Qwen3-8B, exceeding same-backbone published results of 0.841 and 0.809. The full system reaches 0.872 and 0.827, respectively. The gains persist from 0.6B to 8B, with a leak-free 5-fold result of 0.8707 ± 0.0011 on Llama. The verifier covers 23/25 IFEval constraint families and transfers to RewardBench-v1. These results show that richer supervision can improve reward models without enlarging the backbone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.