Learning The Changes: Transferable Visual Transformations From Image-Pair Compression
Abstract
A pair of images demonstrates a visual change precisely but binds it to the identity, layout, and texture of one scene. Existing systems typically learn to reuse the change from matched demonstration–query examples or represent it in a pretrained semantic space. We ask whether a reusable representation can instead emerge from reconstruction alone. A pair encoder compresses a source–target pair into a short edit code of a few continuous tokens, and a jointly trained diffusion editor reconstructs the target from it, with no text and no example of the edit acting on a second image. The learned codes apply the demonstrated change to new pairings of source images and transformations. Longer codes reconstruct the demonstrated pair more faithfully; shorter codes transfer more accurately and organize more clearly by transformation. The editor also applies two compatible changes at once from the concatenated codes of separate demonstrations, a format it never receives in training. This compositionality emerges from reconstruction as well, once the training targets themselves contain more than one change, as natural-image edits often do. The compressed representation thus retains the transformation itself, enough to carry it to new images and to combine it with another.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.