TailSteer: Trajectory-Tail Steering for Drift-Resistant Multi-turn Image Editing
Abstract
When using most generative image editors to edit an image over multiple turns, minor unintended changes may accumulate, resulting in a corrupted result in the end. This long-term editing drift is normally mitigated through heavy large-scale pretraining or additional modules that may introduce inference overhead. In this work, we introduce TailSteer, a simple post-training framework that improves the multi-turn robustness of pretrained editors without changing their architecture or introducing additional computational overhead. TailSteer follows a simple principle: drift is often introduced in the last few denoising steps. Therefore, it keeps earlier denoising dynamics mostly unchanged and corrects only the tail of the diffusion trajectory to preserve fine-grained details in the source image. These corrections are learned from the editor's own denoising rollouts, using unchanged source content as a preservation reference rather than requiring curated ground-truth editing pairs. Experiments on three backbones (FLUX 1, FLUX 2, and SenseNova families) show that TailSteer can consistently reduce cumulative non-target drift: after TailSteer fine-tuning, the models can normally support 5–10 editing turns without significant errors, while the original models often produce obvious artifacts in fewer than five turns.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.