acceptodds
Under review as a conference paper at ICLR 2027

Right Identities, Wrong Places: Reference Addressing in Multi-Reference Image Generation

Abstract

Multi-reference image generation integrates multiple visual references under textual instructions. When the prompt specifies spatial relations among these references, the model must ensure not only identity fidelity, but also reference addressing: assigning each reference to its designated spatial slot. Existing methods and evaluations, however, largely overlook reference addressing. Through a controlled counterfactual audit, we show that this is a distinct capability gap: models can preserve the requested identities yet fail to execute their prompt-specified reference-to-slot assignment, independent of identity loss, reference count, and physical input-order dependence. We introduce AddrBind, an inference-time executor that aligns an explicit reference-assignment plan with the model's emerging layout and redistributes, rather than amplifies, existing reference-conditioning influence, softly abstaining where the assignment is ambiguous. Across five reference-conditioning configurations and two benchmarks, AddrBind improves slot and exact-assignment accuracy in every configuration, with relative gains of and on our primary CogCanvas backbone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.