acceptodds
Under review as a conference paper at ICLR 2027

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

Abstract

Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by the statistical correlation with human ratings. However, as these metrics are increasingly used as optimization objectives, this alignment by correlation is no longer sufficient: agents may strategize to game the evaluation metric. We study this issue through two complementary notions of alignment. A metric is statistically aligned if it correlates with human ratings, and strategically aligned if it resists strategic perturbations that do not add task-relevant information. We make two contributions. First, we propose test principles for reference-based metrics, consisting of human-rating correlation, degradation sensitivity, and manipulation robustness. These principles evaluate whether a metric agrees with human judgments, penalizes low-effort information loss, and resists strategic score inflation. Second, we develop a unified design framework for mutual-information-based metrics that are designed to resist strategic perturbations. The framework decomposes existing and new metrics into four design choices: information measure, estimation method, text representation, and prediction mechanism. Across peer review, summarization, and question answering tasks, we find that strong human-rating correlation does not imply strategic alignment: LLM-as-a-Judge achieves high correlation but is susceptible to manipulations. In contrast, mutual-information-based metrics substantially improve manipulation robustness. Our framework also uncovers a new metric that achieves the strongest overall robustness in our experiments while remaining competitive on human-rating correlation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.