What Do Labels Let Us Conclude?
Abstract
The same human responses can be evaluated against a single answer or a response distribution. Choosing between these targets can change what counts as a good prediction, and making sense of label disagreement requires evidence about how the data was labelled. We propose a framework for making these evaluation judgements explicit. A measurement contract states the target, and a Mediated-Label Reliability Graph records the production evidence. Analysts use shared rules to determine the supported conclusions. A compiler then combines this record with the requested performance claim to construct an evaluation protocol. Controlled evidence profiles show how removing target or process information changes rule matches and eligible actions while responses stay fixed. In real-data comparisons, changing the target and score changes two model rankings. Applying the complete requirements of twelve performance claims withdraws support for six and narrows one. Matched image and language experiments show a further consequence of target choice: majority reduction increases divergence from the response-distribution target. Together, these studies show how explicit targets and evidence guide model selection, the interpretation of performance gains and label treatment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.