Did the Model Get Better or Just Different? Directional Evaluation of LLM Updates from Unlabeled Probes
Abstract
Model updates can change behavior without revealing whether task reliability improved or degraded. We show that this distinction is structural: symmetric behavioral distances can quantify change magnitude but cannot consistently encode the sign of a reliability change. We introduce the Directional Disagreement Estimator (DDE), a parameter-free estimator that combines the fraction of unlabeled probes on which two model versions disagree with a blind judge’s net preference over those disagreements. Under a simple judge-error model, DDE is attenuated toward zero as judge discriminative accuracy approaches chance. In a pre-specified confirmatory evaluation on GPQA across 22 model-update transitions, our primary endpoint—the correlation of DDE with held-out signed accuracy change—is Spearman ρ=0.600 (95% CI [0.22, 0.87], p=0.0031), and DDE attains AUROC 0.85 for detecting regressions exceeding five accuracy points. A secondary sign-accuracy comparison gives 0.68 against the 0.50 constant-sign baseline on a balanced set. DDE does not outperform a zero-change predictor on absolute error, so we position it as a directional statistic rather than a calibrated estimate of magnitude. Symmetric answer disagreement instead predicts change magnitude but not direction. A judge measured above chance yields useful DDE signal; one measured at chance yields none. Labeled canaries remain stronger, positioning DDE as a screening method that needs no probe labels rather than a replacement for labeled evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.