LatentSwitch: One-Step Text-Guided Image Editing via Switchable Noisy Latents
Abstract
Text-guided image editing aims to modify an input image according to a target prompt while preserving its structure and layout. While diffusion models provide strong generative priors for this task, most prior methods rely on iterative denoising, resulting in high inference cost and unintended changes to unedited regions. To address this, we propose LatentSwitch, a one-step diffusion framework based on the observation that the same latent can produce different semantic contents under different text conditions. Specifically, given a source image and a target prompt, LatentSwitch predicts a shared (switchable) latent that preserves source-specific visual information. Under the source prompt, this latent reconstructs the input image; by switching only the text condition to the target prompt, the pretrained diffusion denoiser directly produces the desired edit in a single step. The latent noise level is adaptively chosen based on the required editing strength, enabling low-noise latents to preserve source details for localized edits and higher-noise latents to support larger semantic changes. Experiments show superior editing quality with substantially lower inference cost than prior image editing methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.