acceptodds
Under review as a conference paper at ICLR 2027

Context-Aware Saliency Control for Diffusion-Based Object Modification

Abstract

Modern diffusion-based inpainting models and image-conditioned adapters provide reference-guided object insertion, but offer limited explicit control over the prominence of generated content. We introduce CASAdapter (Context-Aware Saliency Adapter), a module with 9.7M trainable parameters that modulates an object's CLIP image embedding according to the surrounding scene. Its orthogonal tangent-space update help to preserves the embedding norm after renormalization and bounds its angular displacement. This geometric constraint motivates limiting conditioning drift, without guaranteeing preservation of the generated object's identity. CASAdapter is trained through truncated DDIM sampling using a UniSal-based objective together with semantic and color-gradient regularization. Evaluation on a 300-image MS-COCO validation subset shows positive mean saliency gains under UniSal, DeepGaze IIE, SUM, and both TranSalNet variants at positive modulation settings, and negative mean gains at the tested negative setting. These observations support transfer beyond the training predictor. Larger orthogonal modulation produces higher mean saliency gains alongside lower CLIP similarity. Compared with the reported affine setting, the orthogonal variant offers a more favorable observed trade-off for four of the five predictors at a moderate setting; the strongest tested setting exceeds the affine setting in mean gain under all five predictors while retaining higher CLIP similarity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.