acceptodds
Under review as a conference paper at ICLR 2027

Reference Regulation in In-Context Image Generation for Inverse Problems

Abstract

Pretrained in-context image generation models, which we refer to as image/text-to-image (IT2I) models, accept a reference image alongside a text prompt. Given a reference derived from a degraded measurement, such a model acts as an amortized, measurement-dependent prior that supplies the spatial structure a text prompt leaves unspecified, and data-consistency updates correct it with the exact likelihood. The measurement enters twice, and how much the prior is trusted becomes a design choice. In native IT2I generation from degraded references, we observe early structure preservation, where the observed layout appears at the first step and later steps refine details, and reduced reference influence, where generations that depart from the observed layout respond less to the reference from the start. A reference alone is thus not enough. We propose RefReg, a posterior sampler that regulates reference influence at three levels: constructing a measurement-derived reference, scheduling reference guidance to emphasize structure early and detail later, and suppressing generation-to-text attention during early structure formation. To compare priors on a common footing, we derive an upper bound on the average posterior approximation error of a conditional prior under data-consistency updates. By this criterion, image conditioning lowers the error bound throughout sampling in generations that preserve the observed layout on IT2I backbones. These results make a case for building inverse solvers on IT2I rather than text-only priors. Across linear, nonlinear, and partially blind inverse problems against flow-based solvers, RefReg improves perceptual quality with competitive fidelity and remains robust with as few as eight sampling steps.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.