Bounding the Initial-Noise Density Ratio Keeps Every Image under Reward Alignment
Abstract
Reward fine-tuning raises the score of text-to-image models under a reward model but makes many correct images of a prompt hard to sample: on Stable Diffusion 3.5-Medium, DRaFT raises PickScore by 1.41 while the share of base-model images its outputs cover falls from 94.8% to 34.7%. Penalties on KL divergence, diversity, or update size control averages, which a rare composition barely moves. We instead require every set of images for a prompt to keep its probability within fixed factors of the base model. Our key observation is that a frozen sampler maps each initial noise to one image, so the probability of an image set is the Gaussian mass of its noise preimage, and bounding the density ratio of the initial noise bounds every image set at once. RATIO learns a prompt-conditioned invertible noise transform whose density ratio stays in at every point, for any latent dimension, throughout training: Gaussian-preserving rotations choose directions at no cost, and monotone maps in Gaussian-CDF coordinates move probability under one budget shared by all coordinates. At the same PickScore gain, RATIO keeps 85.7% recall against 69.2–80.6% for seven fine-tuning and noise-space methods, and tests on 5,051 GenEval image events find no violation of its bounds, against 179–386 for the controls. Across three generators and three rewards, the bound fixes how much probability moves, and the reward decides which images receive it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.