TAPER: Token Admission with Priced Early Release for Masked Diffusion Decoding
Abstract
Without a key–value cache, every call of a blockwise masked diffusion decoder predicts all masked positions, yet only the active block may be committed. We study (Token Admission with Priced Early Release), a training-free rule that also commits low-risk positions in a bounded window beyond the block by reusing the active block's price. Released positions never count toward the block's completion quota, and the rule adds no model call, history buffer or calibration table. An exact per-output ledger, , shows why early commits are not saved calls: they mostly replace later within-block batching, so the aggregate call reduction per early commit is only 0.09–0.12. On MATH500, GSM8K, HumanEval and MBPP, windows chosen on these benchmarks cut calls by 6.1% (LLaDA-8B) and 7.9% (Dream-7B) at the same price, with paired accuracy changes of [, ] and [, ] points; the reductions replicate across hardware, prices, lengths and 1,142 development problems. Controls locate the saving: with end-of-sequence tail completion applied to both policies, release still removes about 6% of calls, and at equal admission a confidence-and-history gate with AHD's released thresholds saves little because of its history condition, while confidence alone recovers most of the saving, which limits claims for the price gate itself. Release is a second knob rather than a better operating curve: a fixed price does not fix an accuracy target, we establish neither accuracy non-inferiority nor a reproducible advantage over retuning the admission price, and at length 128 accuracy point estimates fall by up to about one point.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.