From Scores to Claims: A Validity Argument Framework for Large Language Model Evaluation
Abstract
Benchmark scores are widely used as evidence of large language model capabilities, yet their meaning depends on the tasks, protocols, and scoring procedures that produce them. We introduce Score-to-Claim Auditing (S2C), a validity-argument framework grounded in language testing and educational measurement. Given an intended interpretation, S2C examines annotated item demands, changes in scores and rankings across protocols, and automated ratings through expert comparisons and controlled tests. Benchmark Validity Cards synthesize the evidence into supported interpretations and conditions of use. We apply S2C to multiple-choice tasks, mathematical word problems with numerical answers, and automated summary evaluation. In multiple-choice tasks, annotations reveal different demand mixtures beneath a shared response format. Small score changes across protocols can alter model rankings. Among 18 equal-accuracy comparisons, 12 contain offsetting item-level gains and losses. Construct-linked analysis also identifies opposing net changes across demand groups. In summarization, positive correlations with expert ratings coexist with numerical discrepancies. On the same summaries, the judges' point-estimate ordering by consistency correlation with experts reverses between summary-level and system-level aggregation. Controlled source reversals produce substantially larger consistency-rating changes than meaning-preserving expression changes or identical-input repeats, supporting a factual-consistency interpretation under the tested rubric and materials. These findings motivate validity profiles alongside benchmark scores, making the evidence and conditions behind capability and quality claims explicit. Analysis code and numerical data are available at https://anonymous.4open.science/r/s2c-review-E603/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.