Mask-On: Compact Activation-Calibrated Merging for Autoregressive and Diffusion Language Models
Abstract
Autoregressive (AR) and diffusion language models offer complementary generation behaviors, but retaining both pretrained endpoints requires storing two full parameter sets. We propose Mask-On, a gradient-free method that represents both modes using an unchanged AR backbone and a compact AR-to-diffusion update. Building on support/sign encoding, Mask-On compresses AR-to-diffusion weight changes into ternary masks and output-channel scales, calibrated by matching the diffusion endpoint's linear-layer outputs on collected activations. Closed-form scale fitting reduces the search to a single pruning parameter without gradient-based optimization or downstream task-specific tuning. Across four AR–diffusion model pairs and four benchmarks spanning diverse domains, we demonstrate that direct parameter merging does not consistently preserve both generation modes, whereas compact mode-specific updates support more consistent diffusion retention alongside an unchanged AR checkpoint. Mask-On achieves the highest average diffusion-mode accuracy among the compared merging methods on all four models, with an update substantially smaller than a full diffusion checkpoint. Experiments on Fast dLLM v2 show similar quality–throughput trend to the diffusion endpoint across six decoding configurations. A simple training-free experiment using the original AR checkpoint further improves diffusion-draft accuracy by up to 3.26 percentage points on GSM8K, illustrating how access to both modes can support complementary generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.