Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
Abstract
Hardware-agnostic strategies for accelerating text-to-image diffusion, such as timestep distillation and feature caching, can reduce inference time without custom kernels or system-level optimization. Among them, multi-resolution generation strategies have recently received broad attention, achieving more than 5x speedup without training. However, the design of upsampling in the latent space, together with the selective modification of partial regions, causes these methods to exhibit noticeable blurring or artifacts. To this end, we propose MrFlow, a training-free multi-resolution acceleration strategy for pretrained flow-matching models built upon a staged low-to-high-resolution pipeline. MrFlow first rapidly generates the main structure at low resolution, then performs super-resolution in the pixel space using a lightweight pretrained GAN-based model, subsequently injects low-strength noise to enable high-frequency resampling, and finally refines the details at high resolution in one step while preserving structure. Quantitative and qualitative results on FLUX.1-dev and Qwen-Image show that MrFlow exploits quadratic token reduction and reduced step requirement of low-resolution sampling to achieve 10x end-to-end acceleration, while keeping the OneIG-En gap on Qwen-Image within 1%, significantly exceeding other training-free acceleration strategies, and requiring no training or runtime dynamic identification. MrFlow can also be directly combined orthogonally with pre-trained timestep distillation strategies, achieving even higher generation acceleration of up to 25x.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.