Timeout Is Not a Negative Label: Learning from Budget-Censored Verification
Abstract
A training timeout can leave a program's outcome unknown at the verification limit that defines task success. We study randomized verification allocation that preserves the policy gradient of this reference objective. Building on inverse weighting and randomized truncation, we derive the allocation criterion for the smooth improvement bound of an ordinary policy minibatch. Finite batching adds a squared-mean-gradient term to the usual cost–second-moment product, which can change the optimal allocation. Under an unknown trace law and a fixed observation kernel, universally unbiased conditional allocation feedback requires two independent records for the usual product and, in general, three for the finite-batch criterion. Distinct-record estimators attain these orders using only censored observations. In a high-signal control, the optimal verification probability changes from one to approximately , reducing the finite-batch objective by . Censored feedback learns this allocation, although a plug-in estimator is more efficient in the stationary control. A separate executed-program study requires delayed verification for optimal program choice on half the held-out prompts. Negative timeout labels exclude those choices. Checkpoint analysis and exact update moments clarify when allocation improvements have an efficiency interpretation. These results characterize observable allocation feedback without implying a general training-speed advantage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.