Agreement Is Not Coverage: What Evaluation Criteria Register of a Training Change
Abstract
Criteria that monitor optimisation are usually chosen because they agree with human judgements, or with one another, on existing responses. Such agreement does not show whether a criterion registers the changes that training makes. We call this difference the agreement–coverage gap and study it with linear criterion readouts on a shared representation, where the part of a training direction that a readout can register is its overlap with that readout and can be computed before training. Training Qwen2.5-1.5B with DPO along an unmodified reward-model direction already moves eleven FLASK criteria unevenly, in the order of their overlap (Spearman ρ = 0.71). In a controlled test, we remove from a driver criterion's direction only the component shared with a checker criterion and compare two DPO arms at matched KL. Across five pairs whose readouts rank 65%–80% of response pairs alike, the checker's detection of the policy change falls in every pair (aggregate +0.063, 95% CI [0.060, 0.067]) while the driver keeps 79%–99% of its above-chance signal. An equal-length placebo removal has a smaller effect in all three pairs where it could be matched, and pre-training overlap orders the responses of the other ten criteria (ρ = 0.81). The pattern holds at 0.5B and 7B and with a mis-specified checker readout. On generated text, an independent judge sees both criteria move, so our claims concern criterion readouts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.