acceptodds
Under review as a conference paper at ICLR 2027

FreEdit: One-Step Source Memory for Mask-Guided Image Editing

Abstract

Text-to-image (T2I) models can also edit real images according to text prompts. Training-free methods often rely on inversion to preserve source content, requiring repeated evaluations at multiple noise levels and full-image denoising during generation. This is particularly costly for local editing, where the edit region is known in advance and the remaining content should stay unchanged. We present FreEdit (Free-inversion Editing), a training-free and highly efficient method that eliminates multi-step inversion and global denoising for local image editing. FreEdit adds noise to the encoded source at a single intermediate timestep and extracts its attention projections with only one network evaluation. During target denoising, only masked image tokens are updated, attending to stored source keys and values outside the mask for contextual guidance. The generated foreground is combined with the clean source latent, preserving untouched content by construction. For color and style edits, optional QK-CFG further improves source-structure preservation while keeping foreground appearance target-driven. Dedicated appearance evaluations demonstrate strong color fidelity, structural preservation, style alignment, and aesthetic quality. On FLUX.1-dev, FreEdit achieves a 3.58× speedup on the same hardware while improving all six reported mean metrics. Results across Klein, Qwen, and SD3.5 further demonstrate its consistent effectiveness across diverse generative backbones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.