TILT: MODEL-INTRINSIC REWARD ALIGNMENT FOR COMPOSITIONAL DIFFUSION
Abstract
Consider conditional generation where is a prompt composed of multiple concepts . Diffusion models often struggle with compositional prompts, producing samples in which some concepts dominate while others are missing or weakly represented. Prior work attributes these failures to *mode collision*, where single-concept modes of overlap with modes of the joint . To seek out collision-free modes of , or *pure modes*, corrector-based approaches have attempted to suppress collisions at intermediate diffusion times. However, local corrections are often heuristic and do not necessarily steer the generation to a *pure mode* in the final data space. Derived from a principled formulation, we present TILT (**T**est-time model-**I**ntrinsic reward a**L**ignment via **T**ilting), a training-free framework that poses eventual pure mode sampling as a reward for intermediate-time alignment. This reward offers valuable advantages: (1) it is intrinsic to the model, hence external reward models need not be trained by modality-specific datasets, (2) it yields a closed-form target under a variational approximation, which makes it realizable through standard diffusion sampling, and (3) it is interpretable, hence amenable to preference-based modifications. Project page: [https://tinyurl.com/yr5j3bwe](https://tinyurl.com/yr5j3bwe)
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.