PaRDi: Learning to Preserve and Revise with Discrete Diffusion
Abstract
Existing discrete diffusion methods for image editing often encode the source as a reference sequence and generate a separate target from a fully masked canvas—a paradigm we call reference-based regeneration. This forces even content meant to remain unchanged to be generated again, potentially introducing unwanted drift, such as background changes during foreground edits or facial changes during clothing edits. We introduce PaRDi, a discrete diffusion framework that starts from the source canvas and models preservation and repeated revision within the same dynamics. Each token can retain its value or pass through <MASK> before taking a new value, and newly generated tokens remain eligible for further revision. These asynchronous updates produce partially revised canvases in which preserved source tokens and completed revisions jointly guide subsequent generation, supporting both local and image-wide edits. Endpoint-conditioned bridges, combined with constant paths at unchanged positions, analytically provide training-state distributions and transition targets, yielding training canvases and transition-matching objectives without hand-designed remasking trajectories. PaRDi achieves 6.98 on GEdit-EN, 3.95 on ImgEdit, and 92.27% background retention on PIE-Bench. Extending in-place editing to visual reasoning, PaRDi attains 87.29% average accuracy, outperforming DiffThinker's 75.36%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.