acceptodds
Under review as a conference paper at ICLR 2027

Solve One Problem with Another: When Linear Attention Meets Masked Diffusion

Abstract

Long-context language modeling needs a global view of the answer and economical token-level computation. We present LaLanD (Linear Attention Language Diffusion), which unifies discrete diffusion with linear attention: diffusion organizes global answer refinement across rounds, while a bidirectional RWKV-7 denoiser integrates the full context to supply token predictions at every unresolved position. We connect the denoising objective to linear-cost sequence mixing and derive a cost–error criterion that states when an additional denoising round improves the generated distribution under fixed-partition stochastic sampling; Lean-verified proofs establish nested refinement and a training-coverage bound, and an enumerable task confirms the predicted regimes. Following this theory, we build a 3.3B model from RWKV-7 World v3 data with 30B tokens resampled from its own pretraining corpus. It raises the 20-task LongBench mean of its initialization from 11.8 to 19.5 with gains on 18 of 20 tasks, reaches a 68.3% ten-task average, and improves RACE and 4k bAbILong QA1, while retaining a fixed-size recurrent state and mixing for context length , answer length , and rounds, and decoding all answer positions in parallel within each round.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.