acceptodds
Under review as a conference paper at ICLR 2027

Lossless, Yet Not Identical: Execution-Level Numerical Bifurcations in Speculative Decoding

Abstract

Speculative decoding is lossless at the level of the decoding rule, but this guarantee does not require different finite-precision executions to produce identical decisions. We study this distinction in greedy speculative decoding and focus on an attribution problem: when speculative and autoregressive generation disagree, does the disagreement arise specifically under speculative verification, or is the same state already sensitive to the autoregressive execution path? We address this question with a three-way diagnostic that compares incremental autoregressive execution (), full-prefix autoregressive execution (), and speculative verification () at the same token prefix. We treat as verifier-specific disagreement relative to these two matched autoregressive controls. Among 555 reconstructed first-divergence events in Qwen3, 166 satisfy this criterion, while 389 are already autoregressive-path-sensitive. Across Qwen3, SmolLM2, and Pythia, all 1,850 reconstructed BF16 divergence events collapse to under backend-matched FP32 replay, which we use only as a diagnostic control. The observed bifurcations are strongly enriched near low-margin decisions, but their prevalence also varies with execution shape and backend. The results therefore do not support a universal margin threshold or block-size law. A targeted replication on A100/Ampere again identifies the same event class at both tested block sizes, and all 145 reconstructed events collapse under matched FP32 replay. The verifier-specific fraction is nevertheless lower than on RTX4090/Ada, showing that the event class persists across the two tested GPU architectures while its observed magnitude can differ. We further use Selective Canonical Re-verification as a mechanism-grounded intervention. Changing only the execution realization recovers most lost execution identity in most tested settings, while high-margin residuals expose the limits of margin-only selection and direct GPU timing quantifies the cost of the present implementation. These results separate two notions of exactness: losslessness constrains the decoder, whereas execution identity also depends on how that decoder is realized under finite precision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.