acceptodds
Under review as a conference paper at ICLR 2027

MASK-Local Adapters for Parallel Decoding with Frozen Autoregressive Models

Abstract

Adding multi-token prediction to a pretrained language model requires choosing where to allocate trainable capacity and how to accept predictions made without their actual preceding tokens. We study MASK-local residual adapters that update future placeholder positions while freezing the backbone and vocabulary projection. A no-MASK bypass retains the original sequential path, while real-token cache recommit supports approximate parallel decoding. After three training epochs, 3.22M and 4.21M adapter parameters achieve 57.36% and 61.39% future-position agreement on Qwen2.5-7B and Llama3.1-8B, respectively. On matched Qwen evaluation windows, agreement exceeds a PaSS-style embedding baseline by 26.60 percentage points. In a four-benchmark development screen, Llama's clean sequential path reproduces all 214 reference sequences. Its separate confidence-based parallel mode accepts 1.34–2.09 tokens per step, with substantial accuracy losses on HumanEval and BBH. On a 64-question Qwen GSM8K runtime screen, throughput rises from 45.39 to 54.58 tokens/s while accuracy changes from 58/64 to 55/64. Comparisons of confidence, direct-error supervision, and tree-derived risk characterize approximate acceptance within this parameter-efficient extension.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.