acceptodds
Under review as a conference paper at ICLR 2027

Reference-Image-Guided Relation Editing via Token Bottlenecking and Semantic Shift Alignment

Abstract

Existing relation-centric image editing methods mainly rely on textual or spatial priors, which often fail to capture fine-grained relational configurations such as contact locations and relative poses. Compared with textual or spatial guidance, a reference image provides more fine-grained visual relation cues. Therefore, we focus on reference-image-guided relation editing, which transfers the relation depicted in a single reference image to a source image. However, this task presents two challenges. (i) Appearance information is entangled with relation semantics, making it difficult to isolate transferable relation semantics from the reference image. This issue is particularly pronounced in unified multimodal models, where the target generation process can directly attend to reference visual representations. As a result, the generated image relies more on reference appearance than on the intended relation cues. (ii) Relation editing is prone to unintended relation-irrelevant semantic shifts that progressively accumulate during denoising because such edits require coordinated pose and spatial adjustments, whose representations in generative models are coupled with identity, background, and style. To address these challenges, we propose Relation-Aware Consistency-Preserving Editing (RACE), which (i) introduces Relation-Aware Token Bottleneck Attention (RTBA) with a limited-capacity relation-token bottleneck that induces competition between relation and appearance information to promote relation-appearance disentanglement; and (ii) proposes Cross-Modal Semantic Shift Alignment (CSSA) with a Semantic Shift Predictor (SSP). CSSA guides editing toward the intended relation shift by aligning the editing-induced image semantic shift with the intended relation-text semantic shift. Additionally, we construct a 35K-scale Relation Transfer Dataset (RTD-35K) for training. Extensive experiments demonstrate that RACE achieves superior relation transfer accuracy and source-image consistency, while showing applicability to more challenging settings, including multi-pair relation editing and cross-category relation transfer. Our code will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.