Exact Vocabulary Decoding from Far Memory with Groupwise Residual Certificates
Abstract
An autoregressive language model scores every vocabulary row but returns one token. When the output head resides outside accelerator memory, a dense projection must read the full matrix at each step. We show how to preserve the exact decision while retrieving only rows whose upper bounds can beat one exactly evaluated row. Our decoder keeps grouped integer representations and residual error bounds on the accelerator, while the original bfloat16 (BF16) rows remain in host memory. Integer dot products and outward rounding give intervals that contain the correctly rounded BF16 scores. A tie-aware screening rule selects the rows to retrieve and preserves greedy and fixed-counter Gumbel decisions. With groups of 512 coordinates, the resident representation occupies 51.2% of a dense BF16 output head and requires exact evaluation of only 0.005% to 2.190% of rows on average across four primary models, with no failure in 10,240 decisions. Six-model tests and paired generation paths also show no token mismatch. Request time falls by 21.2% to 77.8% against the same canonical dense decoder. These results show that exact offloaded decoding can use sparse row retrieval in place of full-matrix transfer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.