RetryPrec: Precision as a Retry-Time Resource under an External Verifier
Abstract
A quantized language model is often served next to its BF16 original, and a task with a verifier lets a failed attempt be tried again. After a low-precision failure, should the next attempt stay in low precision or switch? The usual rule ranks the two copies by average accuracy and repeats the winner. However, a failure is a selection event: the problems it leaves open lean toward those on which the quantized copy is weak, and at our reference setting a BF16 attempt solves more problems per second from the second failure on. We show that exchanging one low-precision attempt for a high-precision one changes the solve rate by the mean gap times the fraction of problems still open, plus a covariance between staying unsolved and the precision gap. We propose RetryPrec, which runs one block of low-precision attempts and then one block of high-precision attempts, with the block lengths chosen on held-out problems by an unbiased estimate of the solve rate. One such schedule is optimal among the policies that react only to failures, and we derive how failures move the ratio of the two hazards. On twelve settings at matched wall clock RetryPrec reaches points, against for the fixed low-precision schedule and for the best of sixteen other strategies. A test registered before the experiments shows that mixing precisions adds little once each problem uses its better one, so most of the attainable gain lies in choosing a precision for each problem, which failures partly reveal. When the budget also charges the memory held by both checkpoints, the schedule stays in low precision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.