acceptodds
Under review as a conference paper at ICLR 2027

Did the Model Get Better or Just Different? Directional Evaluation of LLM Updates from Unlabeled Probes

Abstract

Model updates can change behavior without revealing whether task reliability improved or degraded. We show that this distinction is structural: symmetric behavioral distances can quantify change magnitude but cannot consistently encode the sign of a reliability change. We introduce the Directional Disagreement Estimator (DDE), a parameter-free estimator that combines the fraction of unlabeled probes on which two model versions disagree with a blind judge’s net preference over those disagreements. Under a simple judge-error model, DDE is attenuated toward zero as judge discriminative accuracy approaches chance. In a pre-specified confirmatory evaluation on GPQA across 22 model-update transitions, our primary endpoint—the correlation of DDE with held-out signed accuracy change—is Spearman ρ=0.600 (95% CI [0.22, 0.87], p=0.0031), and DDE attains AUROC 0.85 for detecting regressions exceeding five accuracy points. A secondary sign-accuracy comparison gives 0.68 against the 0.50 constant-sign baseline on a balanced set. DDE does not outperform a zero-change predictor on absolute error, so we position it as a directional statistic rather than a calibrated estimate of magnitude. Symmetric answer disagreement instead predicts change magnitude but not direction. A judge measured above chance yields useful DDE signal; one measured at chance yields none. Labeled canaries remain stronger, positioning DDE as a screening method that needs no probe labels rather than a replacement for labeled evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.