acceptodds
Under review as a conference paper at ICLR 2027

Rankings Survive, Scales Drift: Hidden-State Error Estimation Across Post-Training

Abstract

Hidden-state probes offer an efficient way to estimate whether a vision-language model (VLM) will answer correctly. Yet post-training changes both the representations a probe reads and the errors it must predict. We ask whether a probe remains useful as the model evolves, separating error ranking from probability calibration. Our central finding is that a frozen probe can retain useful rankings even when its probability estimates become unreliable. We characterize ranking stability in terms of score perturbations and establish that updated error probabilities are not identifiable from unlabeled scores alone. Under a common score shift that preserves the conditional error relation, we derive a finite-sample calibration guarantee for re-centering a calibrated base probe using only unlabeled inputs. Changes in this relation call for labeled recalibration; deteriorating rankings motivate refitting. We further find that sampled disagreement can supervise transferable rankings without correctness labels, and that hidden-state probes remain informative when reasoning precedes the answer and first-token uncertainty is less predictive. Experiments across VLM post-training methods and answer formats support these findings. These results separate the cost of recalibrating a probe from that of relearning its error signal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.