Internal Reference Guidance via Latent Token Merging for Diffusion Sampling
Abstract
Inference-time diffusion sampling guidance plays a key role in enabling high-performance image synthesis by diffusion models. Existing methods, including Classifier-Free Guidance (CFG), typically maintain two denoising branches throughout the iterative denoising process to guide predictions by extrapolating between the two branches, which amounts to increased overhead in terms of the number of function evaluations (NFE). In this paper, we investigate whether effective guidance signals can be constructed within a single denoising branch. We achieve this by shifting guidance extrapolation from the prediction space to the intermediate latent-token feature space within the diffusion backbone. Specifically, we construct an internal reference within the cross-attention blocks by merging and unmerging latent tokens and perform guidance by extrapolating the original attention features away from this coarsened reference. Extensive experiments on diverse datasets show that the proposed Merging and Unmerging Guidance (MUG) serves as a general-purpose plug-in module that improves both U-Net- (\ie, SDXL) and MMDiT-based (\ie, Flux.1-schnell/dev) diffusion models in image fidelity, prompt alignment, and human preference. Notably, MUG fills an important applicability gap by enabling inference-time guidance for natively single-branch models, particularly guidance-distilled models (\eg, DMD2), for which conventional CFG is not inherently supported. MUG generalizes across different tasks, sampling regimes, and guidance paradigms while complementing existing guidance methods from CFG to recent condition-free techniques.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.