acceptodds
Under review as a conference paper at ICLR 2027

Appearance Pointers: Multimodal Region Control of Diffusion Transformers

Abstract

Creative professionals rarely want a single global description of an image; they want region-level control over the material, identity, and placement of individual objects, which text prompting alone cannot reliably deliver. Diffusion Transformers (DiTs) natively ingest heterogeneous text and image tokens, yet provide no mechanism to decide where and how each token should shape the output. We introduce appearance pointers, compact tokens that route a frozen DiT to the correct appearance cue at the correct region by aligning text or reference-image inputs with user-specified masks. A region correspondence network produces these pointers, and a depth-wise aggregation step fuses many regions into a single canvas, so multiple heterogeneous regions are generated in one denoising pass without inflating the token count. Unlike prior region-control methods — tied to a single modality or one region at a time — AppearancePointers is the first to condition the same region on image and text simultaneously, with one set of weights and only 7.5% added parameters. Across text- and image-conditioned benchmarks, and zero-shot on real-world data, our single model matches or surpasses modality-specific state-of-the-art methods, offering a simple, extensible interface for precise multimodal region control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.