ManiGuide: Training-Free Preference Alignment Across Diffusion Models
Abstract
Text-to-image generation has been advancing fast, yet the quality of a single sample still varies widely: users routinely draw many generations and cherry-pick the best one, since even a strong frozen model rarely produces its own best-case quality on a single try. Inspired by the Heeger–Bergen principle of distributional feature matching, we propose ManiGuide, a training-free inference-time guidance method that pushes a frozen generative model's typical output toward its own upper-bound quality as measured by any given metric, through patch-wise sliced-Wasserstein feature matching. Given a scoring metric, we first rank a pool of the model's own generation samples and use the top-ranked images to construct a reference manifold in deep-feature space for self-distillation. Then in generation, we identify patches whose feature distributions deviate most from the reference and guide them toward it using sliced-Wasserstein optimal transport. Across both pixel and latent image generative models, ManiGuide improves or preserves every metric. ManiGuide also generalizes beyond the base model to personalized models: on DreamBooth subject-driven fine-tunes, it turns each personalized model's own best generations into its guidance reference and improves preference metrics without additional training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.