D-Miner: Extract Alignment Data from Masked Diffusion Language Models
Abstract
Masked diffusion language models (DLMs) have emerged as promising alternatives to autoregressive models for language generation. However, their iterative denoising and bidirectional generation may also facilitate the extraction of alignment data used in post-training. Existing extraction methods often rely on target prefixes or costly sampling and filtering, limiting their applicability and efficiency in this setting. In this work, we propose D-Miner, a novel framework that steers the denoising process of DLMs to efficiently generate alignment records from masks alone. Specifically, a lightweight corrector shifts intermediate hidden representations within the frozen DLM. We use membership signals to select the model’s own generations for hidden-state and token-level supervision. We theoretically establish that an ideal corrector can transform the model’s output distribution into one conditioned on successful recovery. Extensive experiments on 12 target DLMs demonstrate improved recovery coverage under semantic reference matching. We further show that the extracted data provides effective supervision for improving downstream coding performance. Code is available at https://anonymous.4open.science/r/D-miner.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.