acceptodds
Under review as a conference paper at ICLR 2027

Shallow Queries, Mature Values: Depth-Asynchronous Self-Speculation for Looped Transformers

Abstract

Looped Transformers reuse a shared block across recurrent depths, making autoregressive decoding expensive because every generated token requires many sequential recurrent passes. Self-speculative decoders reduce this cost by drafting at an early depth and verifying at full depth, but typically bind draft computation to prefix representations from the same recurrent depth. We find that queries and keys approach their final-depth representations earlier than values, and controlled prefix-channel interventions show that mature values substantially improve shallow draft predictions. Motivated by this asymmetry, we introduce **Depth-Asynchronous Self-Speculation (DAS)**, which decouples the depth of draft computation from the depth of verified-prefix representations it reads. Its **Mature-V** primitive lets shallow queries retrieve full-depth prefix values without additional recurrent computation. We further develop **DAS-Wave**, which combines depth-asynchronous prefix reads with carried parallel refinement, progressive block growth, and an independent full-depth verifier. Across four recurrent-model checkpoints and mathematics and code workloads, DAS-Wave achieves **4.00–6.96×** mean throughput speedup over paired full-depth autoregressive decoding in the same inference stack. These results identify prefix-information depth as an effective design axis for recurrent self-speculation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.