acceptodds
Under review as a conference paper at ICLR 2027

Agreement Is Not Confidence: Trajectory Agreement Degrades as Training Converges

Abstract

When a model's logits were never logged, a cheap substitute for confidence is how often the checkpoints saved during training agree with the final prediction. The same signal drives a published selective-prediction method that weights late disagreement most heavily. We show that this signal carries less information the closer the checkpoints are to the end of training. Two elementary bounds make this precise. The calibration error of trajectory agreement is pinned to within δ_W of the model's error rate, where δ_W is the mean disagreement in the checkpoint window. Its AUROC for error detection is at most one half plus the largest class-conditional probability that some checkpoint in the window disagrees with the final prediction. As the window approaches the final checkpoint, agreement tends to the constant 1: its calibration error tends to the error rate and its AUROC to one half. On 3D visual grounding (ScanRefer and Nr3D, four training runs) and CIFAR-10 with ResNet-18 (three seeds), moving a fixed-size window later raises calibration error and lowers AUROC in 13 of 14 run-by-window-size cells, with bootstrap intervals excluding zero; in the last-checkpoint window, 89–94% of errors receive agreement 1. Temperature-scaled softmax and a three-seed deep ensemble on the same rows are calibrated (ECE 0.031 and 0.041 against 0.228 for agreement, on ScanRefer), so the failure belongs to the construction, not to the model. For the published method, late weighting helps where accuracy keeps improving and hurts where it stalls, and its late limit falls to AUROC 0.51–0.54 in every run.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.