When Are Itemwise Rewards Enough? Finite-Sample Audits for Relational Preferences
Abstract
Rubrics specify what to evaluate; representing pairwise evidence within each criterion remains a separate choice. A prompt–rubric comparison table may remain relational or be summarized by one score per response, imposing Bradley–Terry differences. We make this choice testable at prespecified approximation tolerances. A simultaneous finite-sample audit labels each row scalar-adequate, nonintegrable, or indeterminate, allowing itemwise and relational representations to coexist. For a fixed full-flow estimator, Hodge orthogonality, row-additive squared loss, and Cartesian actions reduce an explicit -action source-risk problem to thresholds: projection pays squared cycle bias and retention pays cyclic estimation MSE. Independent audit and recovery blocks yield regret controlled by threshold-straddling intervals. Component-specific low-rank recovery and entropy-regularized Nash stability then produce policy certificates. In paired Bernoulli experiments, the selector lowers weighted source risk by 38.9% versus forced scalar and 23.4% versus forced relational representation, with no observed unsafe projection relative to the independently calibrated registry oracle. Separately, a continuous odd edgewise transform preserves every score-difference flow exactly when linear; expected-win margins can therefore make scalar source logits nonadditive. In 3,219 human votes across 183 exact Summarize-from-Feedback cells, all prompts remain indeterminate over the examined tolerances. In 3,115 UltraFeedback records, nonlinear scalar-score aggregation produces nonadditive payoffs with measurable policy effects. The resulting method projects only certified rows and retains relational structure when the audit is indeterminate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.