Where to Look, Where to Edit: Enhancing Editability in Inversion-Free Editing
Abstract
Inversion-free methods based on rectified-flow models enable efficient image and video editing by avoiding costly inversion. However, their limited editing strength often leads to poor alignment with target prompts. We find that this limitation results from insufficient attention to target-prompt semantics within editing regions. Selectively increasing this attention can improve target-prompt guidance and enable stronger edits while maintaining structural fidelity. This motivates us to propose Spatially Adaptive Attention Transfer Edit (SAATEdit), a training-free and inversion-free framework for more effective image and video editing. SAATEdit extracts spatial priors from source-prompt-guided attention to localize editing regions and uses these priors to compute bounded spatial gains for targeted attention enhancement. Benefiting from spatial priors and targeted attention enhancement, SAATEdit achieves stronger edits and better alignment with target prompts while maintaining structural fidelity. Experiments on PIE-Bench and FiVE-Bench demonstrate that SAATEdit achieves superior image and video editing performance with minimal inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.