Aligned Textual Scoring Rules
Abstract
Reference-based text scoring systems evaluate a reported text by comparing it with a ground truth text. Existing metrics and language-model judges can be vulnerable to strategic manipulations that increase scores without improving the information conveyed. Conversely, scoring systems designed for strategic robustness may align poorly with scores assigned by human or language-model judges. We develop a framework for learning judge-aligned textual scoring rules within proper-scoring-rule families. Our approach encodes reported and ground truth texts in a shared numerical representation, learns a common linear projection, and fits Brier or Log-Cosh scoring rules to target judge scores. We assume that text representations encode linear properties of beliefs expressed over an underlying semantic space. Under this assumption, truthful reporting maximizes the expected score. When the assumption holds only approximately, we bound the gain from strategic reporting by a multiple of the squared representation error. Experiments on peer grading and question answering evaluate both alignment with target scores and robustness to strategic manipulation. The learned scoring rules improve alignment with judge scores while exhibiting robustness to the evaluated manipulation strategies. These findings support learning within proper-scoring-rule families as an approach to aligning textual evaluation with judge preferences while preserving incentive guarantees under explicit representation assumptions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.