VINS-Distill: Source-Preserving Few-Step Image Editing with Prompt-Rewrite-Augmented Reinforcement Learning
Abstract
Few-step distillation for image editing requires reducing inference cost while maintaining editing capability and preserving fine details of the source image. However, editing teachers may alter content beyond the requested edit, and few-step distillation can inherit and amplify these errors. To address this problem, we introduce VINS-Distill, which uses source-anchored target projection (SATP) to correct unintended changes in the teacher’s distillation targets and prompt-rewrite-augmented reinforcement learning (PRA-RL) to strengthen editing capability. Specifically, SATP replaces teacher denoising targets in non-edited regions with source-derived latent values. PRA-RL further strengthens editing capability by injecting rewritten instructions during denoising while retaining the original instruction for reward evaluation. Together, SATP and PRA-RL enable the four-step editor to preserve source content while maintaining strong editing capability, without requiring rewritten instructions at inference. To comprehensively evaluate the source preservation and editing quality of our method, we introduce VINS-4K-Bench, comprising 3,305 tasks across diverse source images and edit categories to address the scarcity and limited diversity of existing 4K editing benchmarks. Evaluation on VINS-4K-Bench and two established 1K benchmarks shows that our four-step editor consistently improves source preservation and editing quality over existing state-of-the-art few-step baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.