SparseEdit: Stage-Aware Sparse Inference for DiT-Based Image Editing
Abstract
Instruction-based image editing with Diffusion Transformers repeatedly processes generated-latent and reference-image latent tokens at every denoising step, leading to substantial inference cost. We study how to organize sparse update patches over denoising under an explicit per-step update budget. A diagnostic based on velocity differences computed from full forward passes shows a transition from larger and more connected update patches early in denoising to increasingly fragmented masks later. This observation motivates SparseEdit, a training-free sparse inference method centered on stage-dependent spatial selection. Its Patch Selector shifts from priorities averaged over neighborhoods to pointwise priorities as denoising progresses while maintaining an exact update budget. A short dense warmup assigns request-level update and velocity-reuse budgets, and an Execution Controller realizes the selected updates through sparse computation, dense refreshes, and velocity reuse. Across GEdit-Bench and KontextBench, SparseEdit achieves – end-to-end speedups on FLUX.1 Kontext Dev, LongCat-Image-Edit, and Qwen-Image-Edit, with PSNR – dB, SSIM –, and LPIPS – relative to dense outputs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.