acceptodds
Under review as a conference paper at ICLR 2027

SlotWeave: Dependency-Aware Adaptive Slot Scheduling for Diffusion Language Models

Abstract

Masked diffusion large language models (dLLMs) enable parallel token generation, but independently decoding dependent positions can produce incoherent outputs. Slot-based diffusion decoding addresses this issue through autoregressive infilling within fixed-size slots and selecting decodable parallel slots by confidence. However, fixed slot boundaries can separate strongly coupled tokens, while confidence-based selection does not explicitly prevent mutually dependent slots from being decoded together. We propose SlotWeave, a training-free dependency-aware adaptive scheduler for slot-based dLLM decoding. SlotWeave estimates dependencies from attention weights obtained during the decoding forward pass. It first adaptively partitions remaining masked positions into contiguous slots by merging highly coupled neighbors. Then SlotWeave constructs a slot-level dependency graph and greedily selects a non-conflicting set of slots for parallel autoregressive infilling. Experiments on mathematical reasoning and code generation benchmarks show accuracy gains over the fixed slot baseline. For example, on MATH500 at primitive slot size , it improves accuracy from 49.4% to 53.2% with an approximately 10% reduction in throughput. Ablation studies further support the design choices of our method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.