acceptodds
Under review as a conference paper at ICLR 2027

Auditing Early Stopping after Deployment: Sequential Inference with Selective Feedback

Abstract

Changes after deployment can invalidate rules that stop computation or measurement once intermediate results appear sufficient. Early stopping then hides the completed output needed to check its decisions. We study disagreement between early and completed outputs, and show that some violations are statistically indistinguishable from deployments meeting the risk target using the remaining observations alone. Randomised completion restores feedback on selected early stops; under our protocol, it also replaces their proposed outputs at additional cost. We develop sequential inference for this coupling of selective observation and intervention. An information lower bound shows how audit frequency limits detection speed when pre-audit observations reveal no change. A likelihood-ratio alarm matches its inverse audit-frequency dependence under uniform auditing over the analysed risk range, with grid, restart and overshoot penalties. Correcting for selective observation yields anytime-valid upper bounds on average disagreement risk since deployment began under changing distributions. Alarms separately control deployment-wide false-alarm probability when conditional risk always meets its target. Experiments across magnetic resonance imaging, language-model sampling and image classification distinguish observable change from excess disagreement and examine whether returning the reference improves task outcomes. In the primary classifier comparison, adaptive auditing achieves shorter detection delay and lower delivered disagreement at a common normal-operation cost target, with higher post-change expenditure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.