Locate, Mask, and Regenerate: Parser-Guided Self-Repair for Structured Generation with Diffusion Language Models
Abstract
Diffusion language models (DLMs) have attracted growing attention as promising backbones for large language model (LLM)-based agents due to their capabilities for parallel generation and partial revision. To interact successfully with external environments, these agents must reliably generate structured outputs that conform to strict syntactic rules. However, we consistently observe low parsing success rates across a range of structured generation tasks, with further declines as the number of denoising steps decreases. To address this issue, we introduce a new post-generation strategy, namely Parser-Guided Self-Repair (PGSR), which reinterprets the parser feedback as an explicit localization signal for selective refinement. Using the parser-reported error position as an anchor, PGSR remasks a small surrounding span and reconstructs it through DLM infilling. It is designed to perform localized repair without additional training and to generalize across diverse structured output domains in a format-agnostic manner. We conduct comprehensive experiments with LLaDA-8B and Dream-7B across four structured generation tasks: tool calling, code generation, text-to-structured query language (SQL), and molecular generation. Experimental results demonstrate that PGSR consistently improves both structural validity and downstream task performance across different denoising budgets. It is also confirmed that the proposed approach is compatible with existing decoding strategies, yielding further gains when combined with constrained decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.