Align and Recover: Training-Free Guidance for Text-to-Image Diffusion
Abstract
Training-free guidance improves text-to-image synthesis by adjusting the sampling process of a pretrained diffusion model, although stronger prompt adherence can come at the expense of local visual quality. In self-guidance, contrasting original and perturbed predictions provides a sampling direction, but amplifying the entire difference couples structural adjustments with fine-scale changes. To reconcile these effects, we propose Align–Recover Guidance (), which uses the perturbation response for structural alignment and the denoised estimate for appearance refinement. Spectral–Semantic Alignment combines a low-pass prediction contrast with concentrated text attention, while local contrast left unrecovered by this restriction is addressed through Denoised-Content Recovery, which enhances the estimate's high-pass residual. As noise decreases, Trajectory Harmonization progressively relaxes alignment, allowing content recovery to continue under weaker structural guidance. ARG requires no retraining, backpropagation, or external reward model. On SDXL, ARG reduces FID by on COCO-2014 30k and on COCO-2017 5k relative to published SSG.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.