FLASHLight: Preserving Accepted Prefixes in Low-Bit Speculative Decoding
Abstract
Target quantization accelerates speculative verification, making draft execution cost increasingly important. W4A4 drafting reduces this cost, but lower verifier agreement can offset the gain. On DFlash and DSpark, the conditional-acceptance gap to a 16-bit draft widens toward later positions even under basis-consistent rotation-based PTQ (Aligned PTQ). We call this phenomenon quantization-induced suffix decay. We propose FLASHLight, which reframes W4A4 draft quantization as selecting a low-bit representation for the deployed verifier rather than reconstructing the high-precision draft. With target and draft weights frozen, FLASHLight calibrates the quantization-sensitive interface, then selects an independent draft basis using verifier guidance and a prefix-coupled overlap objective for accepted-prefix preservation. On DFlash and DSpark, FLASHLight retains 97.21-97.68% of the greedy-decoding AL of the 16-bit draft paired with the same W4A4 target, improving average AL over Aligned PTQ by approximately 11.8-22.8%. Repeated measurements on a single RTX 4090 show 28.5-29.8% higher mean end-to-end throughput at batch size 1 than that reference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.