HiBlock: Learning Hierarchical Block-Size Policy for Cross-Domain Generalisation in Diffusion Large Language Models
Abstract
Recent advances in diffusion large language models (dLLMs) enable efficient block-based generation in single-domain settings, where blocks have fixed sizes and tokens within each block are denoised jointly. However, directly extending fixed-size block decoding to multi-domain scenarios is challenging because different domains naturally favour different block size strategies and no single size performs best across all domains. Recent adaptive block-size methods attempt to address this by adding heuristic rules that extend block boundaries to certain predefined tokens or newline characters. However, these methods rely on external rules and prompt-specific cues, thereby not enabling dLLMs to explicitly learn adaptive block-size selection ability for cross-domain generalisation. To address this, we propose HiBlock, a Hierarchical Block-size policy learning framework via Group Relative Policy Optimization (GRPO). HiBlock comprises a high-level policy that selects a block-size action and a low-level policy that generates the response under the selected block size, transforming block size from a predefined hyperparameter into a learnable action. During post-training, HiBlock explores diverse block-size actions across GRPO rollouts, while during inference, the learnt policy selects a block size for each input task before response generation. Experiments on 13 benchmarks show its effectiveness in dLLM cross-domain generalisation. Our code is available at https://anonymous.4open.science/r/HiBlock-2026.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.