acceptodds
Under review as a conference paper at ICLR 2027

SpaDiff: Spatial Conditioning Adapter for Open-Vocabulary Segmentation-to-Image

Abstract

When creating visual designs, a highly desired feature is precise spatial control over image generation. Segmentation-mask-to-image (S2I) generation provides such control by binding a textual description, or regional prompt, to each specific image region present in the segmentation mask during generation. While open-vocabulary S2I methods allow arbitrary prompts, they fine-tune the entire base model, so other extensions trained on the original weights no longer combine with them reliably. Adapters, lightweight modules trained on top of a frozen base model, preserve this composability, but existing S2I adapters are restricted to closed vocabularies and built on older base models, transferring poorly to modern Multi-Modal Diffusion Transformers (MM-DiTs). We present , an open-vocabulary S2I adapter for the MM-DiT FLUX that keeps the base model frozen. SpaDiff combines two mechanisms: (1) an adaptive attention masking strategy that binds regional prompts to their corresponding image region by suppressing image and text tokens interactions softly with a learned per-timestep, per-layer, and per-head strength to avoid disrupting the frozen base model; and (2) a mechanism for adhering to exact mask boundaries, explored in two variants, along with a lightweight feature gating mechanism. SpaDiff achieves state-of-the-art performance among adapter-based methods on both the closed-vocabulary COCO-Stuff and open-vocabulary SACap-Eval benchmarks, while approaching the performance of methods that fine-tune the full base model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.