acceptodds
Under review as a conference paper at ICLR 2027

CrossPair: Unlocking Generalist Visual Transformation Learning in Image Editors

Abstract

Language models switch tasks from demonstrations, but vision still often relies on separate systems for restoration, editing, and stylization. A before–after example offers a shared instruction only if the transferable change can be separated from the scene that demonstrates it. We introduce CrossPair, which adapts a pretrained multimodal image editor to transfer diverse transformations from one exemplar pair. Exemplar-before, exemplar-after, and query images take fixed roles in its native VLM and VAE pathways: the VLM extracts operation semantics, while the VAE retains fine appearance and structure. A training-only, change-aware DINOv3 objective aligns feature-change directions across content-distinct scenes, guiding the output to follow the demonstrated relation without an inference-time encoder. We construct Transform-360K by grouping filtered pairs from conditional representations, restoration, editing, virtual try-on, and stylization under shared transformations. On Relation252K-derived seen and held-out transformations, \method improves target agreement and edit-direction similarity over released systems. Ablations show that the semantic and visual exemplar routes are complementary and that DINO relation alignment further strengthens transfer; Transform-360K training also improves aggregate held-out-operation scores.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.