REVE: Region-Aware Velocity Extrapolation for Accelerated Image Editing
Abstract
Instruction-based image editing models produce high-quality edits but often incur substantial inference latency due to repeated evaluations of diffusion transformers. Existing training-free acceleration methods reduce this inference cost by reusing cached velocities or rescaling their magnitudes at skipped timesteps, but these approximations often compromise fine detail. We observe that velocity extrapolation improves fine detail in edit regions but introduces artifacts in non-edit regions, while velocity reuse helps preserve non-edit content. Building on this insight, we propose REgion-aware Velocity Extrapolation (REVE), a training-free framework that approximates velocities at skipped timesteps in a region-aware and temporally adaptive manner. REVE distinguishes edit from non-edit regions by measuring agreement between predicted velocities and reference velocities that reconstruct the reference image. At skipped timesteps, REVE reuses velocities in non-edit regions and extrapolates velocities in edit regions, with a time-dependent weight that progressively strengthens extrapolation as denoising advances from structure formation to detail refinement. Experiments with Step1X-Edit, Boogu-Image-0.1-Edit, and LongCat-Image-Edit on GEdit-Bench and GIE-Bench show that REVE achieves – speedups, matching or exceeding those of the compared training-free acceleration methods. Among these methods, REVE achieves the best semantic consistency, perceptual quality, and detail fidelity, together with the highest pixel-level and perceptual similarity to the unaccelerated results.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.